Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (14)

Search Parameters:
Keywords = SSDLite

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
17 pages, 610 KB  
Article
Task-Specific Detector Adaptation for Edge MOT: Tracking and Deployment Trade-Offs on the NVIDIA Jetson Nano
by Bruna de Vargas Guterres, Juan Pedro de León, Pablo D. Cuña, Víctor Castelli, Silvia Silva da Costa Botelho and Marcelo Rita Pias
J. Imaging 2026, 12(9), 411; https://doi.org/10.3390/jimaging12090411 - 1 Sep 2026
Viewed by 249
Abstract
Although several MOT solutions have been proposed, limited evidence is available regarding how task-specific detector adaptation affects tracking quality and deployment requirements on low-cost edge hardware. These effects have not been jointly evaluated under fixed tracking and deployment conditions. This work evaluates MOT [...] Read more.
Although several MOT solutions have been proposed, limited evidence is available regarding how task-specific detector adaptation affects tracking quality and deployment requirements on low-cost edge hardware. These effects have not been jointly evaluated under fixed tracking and deployment conditions. This work evaluates MOT on a 4 GB NVIDIA Jetson Nano using the MOT17 benchmark. Experiment 1 served as a baseline characterization and detector-selection stage. YOLOv8n, SSDLite320 and Faster R-CNN were evaluated with a fixed OC-SORT configuration under the same tracking framework. Based on the observed trade-offs, YOLOv8 was selected for adaptation. A task-specific fine-tuning stage was subsequently performed for YOLOv8n and YOLOv8s using more than 30,000 annotated pedestrian image records. Experiment 2 constituted the main analysis. Baseline and adapted models were evaluated under the same tracking and deployment protocol. Tracking performance was assessed through MOTA, IDF1, and HOTA. Processing throughput was measured together with system RAM usage. Board power consumption and energy per frame were also measured. In the paired YOLOv8n comparison, MOTA increased from 15.95 to 51.09 after adaptation. Throughput changed from 3.6 to 3.5 FPS. System RAM usage changed from 3.3820 to 3.3925 GB. Average board power changed from 4235 to 4227 mW. Energy per frame increased by approximately 2.7%. The adapted YOLOv8s model achieved a MOTA of 55.88 at 2.5 FPS. None of the evaluated pipelines achieved conventional real-time video throughput. These findings indicate that task-specific adaptation improved tracking performance with limited changes in the measured deployment characteristics within the evaluated configuration. Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
Show Figures

Figure 1

20 pages, 5380 KB  
Article
SAVE: Spectrum-Aided Visual Enhancement for AI-Based Skin Cancer Detection
by Hung-Yi Huang, Yaswanth Nagisetti, Arvind Mukundan, Riya Karmarkar, Sahaya Ashik Libu, Tao-Yuan Liu and Hsiang-Chen Wang
Diagnostics 2026, 16(12), 1864; https://doi.org/10.3390/diagnostics16121864 - 16 Jun 2026
Cited by 1 | Viewed by 603
Abstract
Background/Objectives: The early identification of skin cancer by standard RGB dermoscopy is a clinical difficulty because of the complex visual differences between impacted lesions and healthy tissue. Methods: For the biomedical challenge, a novel approach to signal processing and image reconstruction is introduced [...] Read more.
Background/Objectives: The early identification of skin cancer by standard RGB dermoscopy is a clinical difficulty because of the complex visual differences between impacted lesions and healthy tissue. Methods: For the biomedical challenge, a novel approach to signal processing and image reconstruction is introduced in this study, called the spectrum-aided visual enhancer (SAVE). The proposed SAVE mechanism aims at reconstructing the diagnostically relevant spectral information from the conventional RGB dermoscopic images using the principles of hyperspectral imaging (HSI) and band selection (BS). After quality control and pre-processing, the images in the ISIC2019 dataset were selected, with 865 images that contain basal cell carcinoma (BCC), seborrheic keratosis (SK), and actinic keratosis (AK) lesions. To reduce data leakage, the dataset was split into training, validation, and testing subsets of 70%, 20%, and 10%, respectively. Five supervised deep learning object detection models were trained and tested on the conventional RGB image dataset and on the SAVE-enhanced dataset. Five supervised deep learning object detection models, namely, YOLOv8, YOLOv10, YOLOv11, SSDLite, and SSD, were trained and tested on the conventional RGB image dataset and the SAVE-enhanced dataset. Additional repeated experimental assessments and statistical comparisons were also carried out to evaluate the improvement in performance. Results: The experimental results showed that the SAVE-based pre-processing always yielded better performance in terms of lesion detection than conventional RGB image processing. The SAVE framework for SSD was evaluated and compared with all other evaluated models and was found to be the most successful, with an accuracy of 96%, a precision of 97%, a recall of 96%, and an F1 score of 96%. Conclusions: The results indicate that the proposed SAVE framework could be a promising RGB-compatible spectral enhancement technique for boosting skin cancer detection and computer-aided dermatologic analysis with the aid of AI. Full article
(This article belongs to the Special Issue Artificial Intelligence in Biomedical Signal and Imaging Processing)
Show Figures

Figure 1

21 pages, 6282 KB  
Article
Comparative Evaluation of Deep Learning Object Detectors for Embedded Weed Detection on Resource-Constrained Platforms
by Nurtay Albanbay, Yerik Nugman, Mukhagali Sagyntay, Azamat Mustafa, Ramona Blanes, Algazy Zhauyt, Rustem Kaiyrov and Nurgali Nurgozhayev
Technologies 2026, 14(5), 265; https://doi.org/10.3390/technologies14050265 - 27 Apr 2026
Viewed by 1033
Abstract
Computer vision–based weed detection plays a critical role in agricultural robotics, enabling accurate, selective weeding. These systems operate on resource-constrained embedded platforms, which introduces a significant trade-off between accuracy and efficiency. This study presents a comparative evaluation of six detection models (YOLOv11n, YOLOv11s, [...] Read more.
Computer vision–based weed detection plays a critical role in agricultural robotics, enabling accurate, selective weeding. These systems operate on resource-constrained embedded platforms, which introduces a significant trade-off between accuracy and efficiency. This study presents a comparative evaluation of six detection models (YOLOv11n, YOLOv11s, SSD-Lite, NanoDet, Faster R-CNN, RT-DETR) for agro-robotic applications, measuring precision, recall, mAP@0.5, and runtime on low-power hard-ware. NanoDet achieved the highest detection accuracy (precision 98.6%, recall 94.2%, mAP@0.5 97.7%). YOLOv11s demonstrated similar performance (mAP@0.5: 96.1%) but required more computation. YOLOv11n provides the most favourable balance between accuracy and throughput (mAP@0.5: 94.6%, 207 FPS on a workstation). On Raspberry Pi 5, light models achieved 3–5 FPS. RT-DETR and Faster R-CNN exhibited high latency (3112–6500 ms/frame), which prevents real-time operation. NanoDet excelled in detection, while YOLOv11n provides the best balance between accuracy and efficiency for limited devices. Full article
Show Figures

Figure 1

23 pages, 32193 KB  
Article
Object Detection on Road: Vehicle’s Detection Based on Re-Training Models on NVIDIA-Jetson Platform
by Sleiter Ramos-Sanchez, Jinmi Lezama, Ricardo Yauri and Joyce Zevallos
J. Imaging 2026, 12(1), 20; https://doi.org/10.3390/jimaging12010020 - 1 Jan 2026
Cited by 4 | Viewed by 2606
Abstract
The increasing use of artificial intelligence (AI) and deep learning (DL) techniques has driven advances in vehicle classification and detection applications for embedded devices with deployment constraints due to computational cost and response time. In the case of urban environments with high traffic [...] Read more.
The increasing use of artificial intelligence (AI) and deep learning (DL) techniques has driven advances in vehicle classification and detection applications for embedded devices with deployment constraints due to computational cost and response time. In the case of urban environments with high traffic congestion, such as the city of Lima, it is important to determine the trade-off between model accuracy, type of embedded system, and the dataset used. This study was developed using a methodology adapted from the CRISP-DM approach, which included the acquisition of traffic videos in the city of Lima, their segmentation, and manual labeling. Subsequently, three SSD-based detection models (MobileNetV1-SSD, MobileNetV2-SSD-Lite, and VGG16-SSD) were trained on the NVIDIA Jetson Orin NX 16 GB platform. The results show that the VGG16-SSD model achieved the highest average precision (mAP 90.7%), with a longer training time, while the MobileNetV1-SSD (512×512) model achieved comparable performance (mAP 90.4%) with a shorter time. Additionally, data augmentation through contrast adjustment improved the detection of minority classes such as Tuk-tuk and Motorcycle. The results indicate that, among the evaluated models, MobileNetV1-SSD (512×512) achieved the best balance between accuracy and computational load for its implementation in ADAS embedded systems in congested urban environments. Full article
(This article belongs to the Special Issue Advances in Machine Learning for Computer Vision Applications)
Show Figures

Figure 1

22 pages, 6976 KB  
Article
Re-Parameterization After Pruning: Lightweight Algorithm Based on UAV Remote Sensing Target Detection
by Yang Yang, Pinde Song, Yongchao Wang and Lijia Cao
Sensors 2024, 24(23), 7711; https://doi.org/10.3390/s24237711 - 2 Dec 2024
Cited by 4 | Viewed by 2691
Abstract
Lightweight object detection algorithms play a paramount role in unmanned aerial vehicles (UAVs) remote sensing. However, UAV remote sensing requires target detection algorithms to have higher inference speeds and greater accuracy in detection. At present, most lightweight object detection algorithms have achieved fast [...] Read more.
Lightweight object detection algorithms play a paramount role in unmanned aerial vehicles (UAVs) remote sensing. However, UAV remote sensing requires target detection algorithms to have higher inference speeds and greater accuracy in detection. At present, most lightweight object detection algorithms have achieved fast inference speed, but their detection precision is not satisfactory. Consequently, this paper presents a refined iteration of the lightweight object detection algorithm to address the above issues. The MobileNetV3 based on the efficient channel attention (ECA) module is used as the backbone network of the model. In addition, the focal and efficient intersection over union (FocalEIoU) is used to improve the regression performance of the algorithm and reduce the false-negative rate. Furthermore, the entire model is pruned using the convolution kernel pruning method. After pruning, model parameters and floating-point operations (FLOPs) on VisDrone and DIOR datasets are reduced to 1.2 M and 1.5 M and 6.2 G and 6.5 G, respectively. The pruned model achieves 49 frames per second (FPS) and 44 FPS inference speeds on Jetson AGX Xavier for VisDrone and DIOR datasets, respectively. To fully exploit the performance of the pruned model, a plug-and-play structural re-parameterization fine-tuning method is proposed. The experimental results show that this fine-tuned method improves mAP@0.5 and mAP@0.5:0.95 by 0.4% on the VisDrone dataset and increases mAP@0.5:0.95 by 0.5% on the DIOR dataset. The proposed algorithm outperforms other mainstream lightweight object detection algorithms (except for FLOPs higher than SSDLite and mAP@0.5 Below YOLOv7 Tiny) in terms of parameters, FLOPs, mAP@0.5, and mAP@0.5:0.95. Furthermore, practical validation tests have also demonstrated that the proposed algorithm significantly reduces instances of missed detection and duplicate detection. Full article
(This article belongs to the Section Remote Sensors)
Show Figures

Figure 1

18 pages, 3145 KB  
Article
Puppet Dynasty Recognition System Based on MobileNetV2
by Xiaona Xie, Zeqian Liu, Yuanshuai Wang, Haoyue Fu, Mengqi Liu, Yingqin Zhang and Jinbo Xu
Entropy 2024, 26(8), 645; https://doi.org/10.3390/e26080645 - 29 Jul 2024
Cited by 5 | Viewed by 1828
Abstract
Traditional image classification usually relies on manual feature extraction; however, with the rapid development of artificial intelligence and intelligent vision technology, deep learning models such as CNNs can automatically extract key features from input images to achieve efficient classification. This study focuses on [...] Read more.
Traditional image classification usually relies on manual feature extraction; however, with the rapid development of artificial intelligence and intelligent vision technology, deep learning models such as CNNs can automatically extract key features from input images to achieve efficient classification. This study focuses on the application of lightweight separable convolutional neural networks in domain-specific image classification tasks. In this paper, we discuss how to use the SSDLite object detection algorithm combined with the MobileNetV2 lightweight convolutional architecture for puppet dynasty recognition from images—a novel and challenging task. By constructing a system that combines object detection and image classification, we aimed to solve the problem of automatic puppet dynasty recognition to reduce manual intervention and improve recognition efficiency and accuracy. We hope that this will have significant implications in the fields of cultural protection and art history research. Full article
(This article belongs to the Section Signal and Data Analysis)
Show Figures

Figure 1

22 pages, 3856 KB  
Article
MixMobileNet: A Mixed Mobile Network for Edge Vision Applications
by Yanju Meng, Peng Wu, Jian Feng and Xiaoming Zhang
Electronics 2024, 13(3), 519; https://doi.org/10.3390/electronics13030519 - 26 Jan 2024
Cited by 10 | Viewed by 4880
Abstract
Currently, vision transformers (ViTs) have rivaled comparable performance to convolutional neural networks (CNNs). However, the computational demands of the transformers’ self-attention mechanism pose challenges for their application on edge devices. Therefore, in this study, we propose a lightweight transformer-based network model called MixMobileNet. [...] Read more.
Currently, vision transformers (ViTs) have rivaled comparable performance to convolutional neural networks (CNNs). However, the computational demands of the transformers’ self-attention mechanism pose challenges for their application on edge devices. Therefore, in this study, we propose a lightweight transformer-based network model called MixMobileNet. Similar to the ResNet block, this model only comprises a MixMobile block (MMb), which combines the efficient local inductive bias with the explicit modeling features of a transformer to achieve the fusion of the local–global feature interactions. For local, we propose the local-feature aggregation encoder (LFAE), which incorporates a PC2P (Partial-Conv→PWconv→PWconv) inverted bottleneck structure for residual connectivity. In particular, the kernel and channel scale are adaptive, reducing feature redundancy in adjacent layers and efficiently representing parameters. For global, we propose the global-feature aggregation encoder (GFAE), which employs a pooling strategy and computes the covariance matrix between channels instead of the spatial dimensions, changing the computational complexity from quadratic to linear, and this accelerates the inference of the model. We perform extensive image classification, object detection, and segmentation experiments to validate model performance. Our MixMobileNet-XXS/XS/S achieves 70.6%/75.1%/78.8% top-1 accuracy with 1.5 M/3.2 M/7.3 M parameters and 0.2 G/0.5 G/1.2 G FLOPs on ImageNet-1K, outperforming MobileViT-XXS/XS/S with an improvement of +1.6%↑/+0.4%↑/+0.4%↑ with −38.8%↓/−51.5%↓/−39.8%↓ reduction in FLOPs. In addition, the MixMobileNet-S assembly of SSDLite and DeepLabv3 achieves an accuracy of 28.5 mAP/79.5 mIoU at COCO2017/VOC2012 with lower computation, demonstrating the competitive performance of our lightweight model. Full article
(This article belongs to the Topic Cloud and Edge Computing for Smart Devices)
Show Figures

Figure 1

20 pages, 37823 KB  
Article
A Real-Time Subway Driver Action Sensoring and Detection Based on Lightweight ShuffleNetV2 Network
by Xing Shen and Xiukun Wei
Sensors 2023, 23(23), 9503; https://doi.org/10.3390/s23239503 - 29 Nov 2023
Cited by 6 | Viewed by 2480
Abstract
The driving operations of the subway system are of great significance in ensuring the safety of trains. There are several hand actions defined in the driving instructions that the driver must strictly execute while operating the train. The actions directly indicate whether equipment [...] Read more.
The driving operations of the subway system are of great significance in ensuring the safety of trains. There are several hand actions defined in the driving instructions that the driver must strictly execute while operating the train. The actions directly indicate whether equipment is normally operating. Therefore, it is important to automatically sense the region of the driver and detect the actions of the driver from surveillance cameras to determine whether they are carrying out the corresponding actions correctly or not. In this paper, a lightweight two-stage model for subway driver action sensoring and detection is proposed, consisting of a driver detection network to sense the region of the driver and an action recognition network to recognize the category of an action. The driver detection network adopts the pretrained MobileNetV2-SSDLite. The action recognition network employs an improved ShuffleNetV2, which incorporates a spatial enhanced module (SEM), improved shuffle units (ISUs), and shuffle attention modules (SAMs). SEM is used to enhance the feature maps after convolutional downsampling. ISU introduces a new branch to expand the receptive field of the network. SAM enables the model to focus on important channels and key spatial locations. Experimental results show that the proposed model outperforms 3D MobileNetV1, 3D MobileNetV3, SlowFast, SlowOnly, and SE-STAD models. Furthermore, a subway driver action sensoring and detection system based on a surveillance camera is built, which is composed of a video-reading module, main operation module, and result-displaying module. The system can perform action sensoring and detection from surveillance cameras directly. According to the runtime analysis, the system meets the requirements for real-time detection. Full article
(This article belongs to the Special Issue Deep Learning Technology and Image Sensing)
Show Figures

Figure 1

9 pages, 417 KB  
Article
SSDLiteX: Enhancing SSDLite for Small Object Detection
by Hyeong-Ju Kang
Appl. Sci. 2023, 13(21), 12001; https://doi.org/10.3390/app132112001 - 3 Nov 2023
Cited by 8 | Viewed by 3574
Abstract
Object detection in many real applications requires the capability of detecting small objects in a system with limited resources. Convolutional neural networks (CNNs) show high performance in object detection, but they are not adequate to resource-limited environments. The combination of MobileNet V2 and [...] Read more.
Object detection in many real applications requires the capability of detecting small objects in a system with limited resources. Convolutional neural networks (CNNs) show high performance in object detection, but they are not adequate to resource-limited environments. The combination of MobileNet V2 and SSDLite is one of the common choices in such environments, but it has a problem in detecting small objects. This paper analyzes the structure of SSDLite and proposes variations leading to small object detection improvement. The feature maps with the higher resolution are utilized more, and the base CNN is modified to have more layers in the high resolution. Experiments have been performed for the various configurations and the results show the proposed CNN, SSDLiteX, improves the detection accuracy AP of small objects by 1.5 percent points in the MS COCO data set. Full article
(This article belongs to the Special Issue Computer Vision and Pattern Recognition Based on Deep Learning)
Show Figures

Figure 1

17 pages, 468 KB  
Article
AoCStream: All-on-Chip CNN Accelerator with Stream-Based Line-Buffer Architecture and Accelerator-Aware Pruning
by Hyeong-Ju Kang and Byung-Do Yang
Sensors 2023, 23(19), 8104; https://doi.org/10.3390/s23198104 - 27 Sep 2023
Cited by 7 | Viewed by 3760
Abstract
Convolutional neural networks (CNNs) play a crucial role in many EdgeAI and TinyML applications, but their implementation usually requires external memory, which degrades the feasibility of such resource-hungry environments. To solve this problem, this paper proposes memory-reduction methods at the algorithm and architecture [...] Read more.
Convolutional neural networks (CNNs) play a crucial role in many EdgeAI and TinyML applications, but their implementation usually requires external memory, which degrades the feasibility of such resource-hungry environments. To solve this problem, this paper proposes memory-reduction methods at the algorithm and architecture level, implementing a reasonable-performance CNN with the on-chip memory of a practical device. At the algorithm level, accelerator-aware pruning is adopted to reduce the weight memory amount. For activation memory reduction, a stream-based line-buffer architecture is proposed. In the proposed architecture, each layer is implemented by a dedicated block, and the layer blocks operate in a pipelined way. Each block has a line buffer to store a few rows of input data instead of a frame buffer to store the whole feature map, reducing intermediate data-storage size. The experimental results show that the object-detection CNNs of MobileNetV1/V2 and an SSDLite variant, widely used in TinyML applications, can be implemented even on a low-end FPGA without external memory. Full article
(This article belongs to the Special Issue The Rise of EdgeAI and TinyML for the Next-Generation IoT)
Show Figures

Figure 1

17 pages, 22048 KB  
Article
Underwater Accompanying Robot Based on SSDLite Gesture Recognition
by Tingzhuang Liu, Yi Zhu, Kefei Wu and Fei Yuan
Appl. Sci. 2022, 12(18), 9131; https://doi.org/10.3390/app12189131 - 11 Sep 2022
Cited by 10 | Viewed by 3501
Abstract
Underwater robots are often used in marine exploration and development to assist divers in underwater tasks. However, the underwater robots on the market have some problems, such as only a single function of object detection or tracking, the use of traditional algorithms with [...] Read more.
Underwater robots are often used in marine exploration and development to assist divers in underwater tasks. However, the underwater robots on the market have some problems, such as only a single function of object detection or tracking, the use of traditional algorithms with low accuracy and robustness, and the lack of effective interaction with divers. To this end, we designed a type of gesture recognition based on interaction, using person tracking as an auxiliary means for an underwater accompanying robot (UAR). We train and test the SSDLite detection algorithm using the self-labeled underwater datasets, and combine the kernelized correlation filters (KCF) tracking algorithm with the “Active Control” target tracking rule to continuously track the underwater human body. Our experiments show that the use of underwater datasets and target tracking can effectively improve gesture recognition accuracy by 40–105%. In the outfield experiment, the performance of the algorithm was good. It achieved target tracking and gesture recognition at 29.4 FPS on Jetson Xavier NX, and the UAR made corresponding actions according to the diver gesture command. Full article
(This article belongs to the Special Issue Underwater Robot)
Show Figures

Figure 1

15 pages, 3932 KB  
Article
A Real-Time FPGA Accelerator Based on Winograd Algorithm for Underwater Object Detection
by Liangwei Cai, Ceng Wang and Yuan Xu
Electronics 2021, 10(23), 2889; https://doi.org/10.3390/electronics10232889 - 23 Nov 2021
Cited by 14 | Viewed by 5074
Abstract
Real-time object detection is a challenging but crucial task for autonomous underwater vehicles because of the complex underwater imaging environment. Resulted by suspended particles scattering and wavelength-dependent light attenuation, underwater images are always hazy and color-distorted. To overcome the difficulties caused by these [...] Read more.
Real-time object detection is a challenging but crucial task for autonomous underwater vehicles because of the complex underwater imaging environment. Resulted by suspended particles scattering and wavelength-dependent light attenuation, underwater images are always hazy and color-distorted. To overcome the difficulties caused by these problems to underwater object detection, an end-to-end CNN network combined U-Net and MobileNetV3-SSDLite is proposed. Furthermore, the FPGA implementation of various convolution in the proposed network is optimized based on the Winograd algorithm. An efficient upsampling engine is presented, and the FPGA implementation of squeeze-and-excitation module in MobileNetV3 is optimized. The accelerator is implemented on a Zynq XC7Z045 device running at 150 MHz and achieves 23.68 frames per second (fps) and 33.14 fps when using MobileNetV3-Large and MobileNetV3-Small as the feature extractor. Compared to CPU, our accelerator achieves 7.5×–8.7× speedup and 52×–60× energy efficiency. Full article
(This article belongs to the Special Issue Recent FPGA Architectures and Applications)
Show Figures

Figure 1

24 pages, 11300 KB  
Article
Visible and Thermal Image-Based Trunk Detection with Deep Learning for Forestry Mobile Robotics
by Daniel Queirós da Silva, Filipe Neves dos Santos, Armando Jorge Sousa and Vítor Filipe
J. Imaging 2021, 7(9), 176; https://doi.org/10.3390/jimaging7090176 - 3 Sep 2021
Cited by 41 | Viewed by 6104
Abstract
Mobile robotics in forests is currently a hugely important topic due to the recurring appearance of forest wildfires. Thus, in-site management of forest inventory and biomass is required. To tackle this issue, this work presents a study on detection at the ground level [...] Read more.
Mobile robotics in forests is currently a hugely important topic due to the recurring appearance of forest wildfires. Thus, in-site management of forest inventory and biomass is required. To tackle this issue, this work presents a study on detection at the ground level of forest tree trunks in visible and thermal images using deep learning-based object detection methods. For this purpose, a forestry dataset composed of 2895 images was built and made publicly available. Using this dataset, five models were trained and benchmarked to detect the tree trunks. The selected models were SSD MobileNetV2, SSD Inception-v2, SSD ResNet50, SSDLite MobileDet and YOLOv4 Tiny. Promising results were obtained; for instance, YOLOv4 Tiny was the best model that achieved the highest AP (90%) and F1 score (89%). The inference time was also evaluated, for these models, on CPU and GPU. The results showed that YOLOv4 Tiny was the fastest detector running on GPU (8 ms). This work will enhance the development of vision perception systems for smarter forestry robots. Full article
Show Figures

Figure 1

36 pages, 5671 KB  
Article
UAV Landing Using Computer Vision Techniques for Human Detection
by David Safadinho, João Ramos, Roberto Ribeiro, Vítor Filipe, João Barroso and António Pereira
Sensors 2020, 20(3), 613; https://doi.org/10.3390/s20030613 - 22 Jan 2020
Cited by 44 | Viewed by 9421
Abstract
The capability of drones to perform autonomous missions has led retail companies to use them for deliveries, saving time and human resources. In these services, the delivery depends on the Global Positioning System (GPS) to define an approximate landing point. However, the landscape [...] Read more.
The capability of drones to perform autonomous missions has led retail companies to use them for deliveries, saving time and human resources. In these services, the delivery depends on the Global Positioning System (GPS) to define an approximate landing point. However, the landscape can interfere with the satellite signal (e.g., tall buildings), reducing the accuracy of this approach. Changes in the environment can also invalidate the security of a previously defined landing site (e.g., irregular terrain, swimming pool). Therefore, the main goal of this work is to improve the process of goods delivery using drones, focusing on the detection of the potential receiver. We developed a solution that has been improved along its iterative assessment composed of five test scenarios. The built prototype complements the GPS through Computer Vision (CV) algorithms, based on Convolutional Neural Networks (CNN), running in a Raspberry Pi 3 with a Pi NoIR Camera (i.e., No InfraRed—without infrared filter). The experiments were performed with the models Single Shot Detector (SSD) MobileNet-V2, and SSDLite-MobileNet-V2. The best results were obtained in the afternoon, with the SSDLite architecture, for distances and heights between 2.5–10 m, with recalls from 59%–76%. The results confirm that a low computing power and cost-effective system can perform aerial human detection, estimating the landing position without an additional visual marker. Full article
(This article belongs to the Special Issue Architectures and Platforms for Smart and Sustainable Cities)
Show Figures

Figure 1

Back to TopTop