1. Introduction
Small-sized Unmanned Aerial Vehicles (UAVs) have been rapidly proliferating across commercial, industrial, and individual application domains in line with ongoing technological advancements. This widespread adoption has also introduced a range of security threats arising from the potential malicious use of these systems. In particular, the increasing security requirements for the protection of critical infrastructure, military facilities, and private properties have made the need for reliable, fast, and cost-effective UAV detection solutions increasingly important. The literature indicates that optical imaging and radar-based detection systems suffer from degraded sensing performance under environmental disturbances such as fog, rain, smoke, and low-light conditions. This limitation highlights the need for the development of alternative detection paradigms that can complement optical and radar systems to achieve more reliable results [
1]. In non-line-of-sight (NLOS) scenarios, as well as in cases involving mini/micro UAVs with low radar cross-section (RCS), target detection performance is significantly reduced due to weak signal returns and low signal-to-noise ratio (SNR) conditions [
1,
2]. These constraints necessitate the development of alternative and complementary detection paradigms in order to achieve more reliable and effective outcomes.
Radio Frequency (RF)-based detection has been reported in the literature as an alternative approach to optical and radar-based methods, relying on the passive monitoring of electromagnetic emissions generated by UAVs during their command-and-control communications [
3,
4]. This approach offers several advantages, including the absence of active illumination requirements, the ability to operate independently of line-of-sight conditions, and relatively low hardware cost. However, the high level of noise and interference present in RF environments poses significant technical challenges for signal classification tasks [
5]. In recent years, deep learning (DL) methods have been widely adopted in RF-based UAV detection systems due to their strong performance in accurately classifying complex and noisy signal environments [
6,
7,
8].
The primary objective of this study is to develop a methodology for the autonomous detection of UAVs using RF signals and DL techniques, enabling high-accuracy classification of a total of seven classes, including six different consumer-grade drone models and noise data. Within the scope of this research, a multimodal classification framework is implemented using the “Noisy Drone RF Signal Classification” dataset, which contains complex spectral characteristics of UAV signals. The proposed approach incorporates both raw IQ signals and spectrogram-based representations. Furthermore, the study conducts a detailed analysis of hybrid Convolutional Neural Network (CNN)-based architectures in noisy RF environments, with the aim of developing a low-cost, effective, and scalable RF-based detection and classification framework suitable for security applications.
2. Related Works
RF-based UAV detection and classification has attracted increasing research attention in the field of security systems in recent years. Studies in this domain exhibit substantial differences in terms of signal representation, model architecture, and targeted application scenarios. When the existing literature is examined, it is observed that research efforts span a broad spectrum ranging from dataset development to DL-based classification approaches. For instance, Al-Sa’d et al. (2019) [
9] developed the open-source DroneRF dataset, which contains RF signals from various UAVs operating under different flight modes. In that study, the applicability of the dataset was evaluated using deep neural networks. The findings revealed that the average accuracy decreased from 99.7% for 2 classes to 46.8% for 10 classes as the number of classes increased, highlighting the inherent difficulty of multi-class UAV identification problems [
9]. Similarly, Allahham et al. (2020) [
8,
9] performed UAV detection and state classification on the same dataset using a multi-channel 1D-CNN architecture. The reported results further confirmed the performance degradation with increasing class complexity, achieving 100% accuracy for 2 classes, 94.6% for 4 classes, and 87.4% for 10 classes. This observed decline in accuracy with respect to class number is consistent with the 87.42% performance obtained in the present study on a 7-class noisy dataset.
Studies adopting traditional machine learning methods have also contributed significantly to the field. Ezuma et al. (2020) [
10] classified 15 different UAV controllers using RF fingerprinting in the presence of Wi-Fi and Bluetooth interference signals. In that work, an accuracy of 98.13% was achieved using a k-nearest neighbors (kNN) classifier under 25 dB SNR conditions. However, the study also reported a significant performance degradation when the SNR dropped below 10 dB. Yang et al. (2021) [
11] proposed a hybrid approach combining a feature engineering generator (FEG) with a multi-channel deep neural network (MC-DNN), achieving 98.4% accuracy in a ten-class classification task. Nevertheless, the proposed method relies on manual preprocessing steps such as signal truncation and moving average filtering. In contrast to these limitations, the present study processes raw IQ and spectrogram data in an end-to-end manner without any manual intervention.
DL-based approaches, in particular, have demonstrated notable performance advantages in noisy environments. Glüge et al. (2023) [
6] compared 1D IQ and 2D spectrogram representations using CNN architectures for UAV detection from RF signals and showed that spectrogram-based approaches significantly outperform IQ representations under low-SNR conditions (−12 dB). This finding constitutes the primary motivation for integrating both representations into a unified multimodal architecture in the present study. Alam et al. (2023) [
7] proposed an end-to-end DL architecture capable of learning complex signal patterns through residual blocks, achieving 97.53% overall accuracy and a 0.37 ms inference time over the CardRF dataset across an SNR range of 0 dB to 30 dB. Similarly, Zhao et al. (2024) [
12] proposed a spectrogram-based DL model for detecting UAVs employing frequency-hopping spread spectrum signals, achieving 90.9% classification accuracy across different UAV types, thereby confirming the effectiveness of spectrogram representations in complex RF environments. Zhang et al. (2023) [
13] further enhanced neural network robustness and generalization capability in complex electromagnetic environments through low-cost data augmentation and spectrogram segmentation techniques. The reported findings collectively support the effectiveness of spectrogram-based representations in noisy RF scenarios.
With respect to model robustness and modality fusion, several notable studies have also been reported. Elyousseph and Altamimi (2024) [
14] reported that image-based CNN classifiers exhibit approximately 40% higher robustness compared to coefficient-based classifiers under low-SNR conditions. This result indicates the superior generalization capability of CNN architectures to channel conditions not seen during training. Frid et al. (2024) [
15], on the other hand, achieved 91% classification accuracy at −10 dB SNR using a multimodal approach based on the fusion of RF and acoustic signals, demonstrating that combining different modalities improves detection performance in low-SNR scenarios.
When the reviewed studies are collectively considered, it becomes evident that the majority of the existing literature either relies on a single signal representation (either IQ or spectrogram) or is evaluated under clean, controlled environments. To address this gap, the present study develops a multimodal hybrid CNN architecture that simultaneously processes raw IQ signals and spectrogram data, and evaluates the model extensively on the “Noisy Drone RF Signal Classification” dataset, which reflects real-world noise conditions.
3. Method
In this section, the dataset used for UAV signal detection and classification, the applied preprocessing steps, and the architecture of the developed multimodal DL model are described in detail. The study is fundamentally based on the simultaneous analysis of raw RF signals in both the time and frequency domains, and the processing of these representations through hybrid CNN architectures.
3.1. Dataset Details
In this study, the UAV RF signal dataset made publicly available by Glüge et al. [
6] is utilized. The dataset consists of RF signals belonging to six different UAV/controller systems—namely DJI (SZ DJI Technology Co., Ltd., Shenzhen, China), FutabaT14 (Futaba Corporation, Mobara, Japan), FutabaT7 (Futaba Corporation, Mobara, Japan), Graupner (Graupner GmbH & Co. KG, Kirchheim/Teck, Germany), Taranis (FrSky Electronic Co., Ltd., Wuxi, China), and Turnigy (HobbyKing, Hong Kong, China)—as well as a noise class, resulting in a total of seven classes. Each sample corresponds to a non-overlapping signal segment of 16,384 points, representing approximately 1.2 ms of data at a 14 MHz sampling rate.
After normalization, the signals are mixed with laboratory-generated noise (Labnoise: Bluetooth, Wi-Fi, and amplifier noise) and Gaussian noise in a 50%/50% ratio. The noise class is synthetically constructed using all possible combinations of these two noise types (Labnoise + Labnoise, Labnoise + Gaussian, Gaussian + Labnoise, and Gaussian + Gaussian).
The SNR distribution of the dataset is designed in a controlled manner. For each class, SNR values are uniformly distributed in the range of −20 dB to +30 dB with 2 dB increments, resulting in approximately 3792–3800 samples per SNR level. Class balance within the dataset is maintained through this design, eliminating the need for additional balancing strategies during model training. This setup enables the evaluation of the model’s generalization capability across a wide SNR range, with performance analyzed at 5 dB intervals within this range.
3.2. Data Preprocessing and Experimental Setup
For training, validation, and testing, the dataset is randomly split into 75% training, 15% validation, and 10% independent test sets. During feature extraction, raw signal vectors are transformed into spectrogram images in the time–frequency domain. To improve convergence speed during training, instead of applying heavy preprocessing or manual normalization steps, the dynamic normalization capabilities of Batch Normalization layers within the model architecture are utilized.
The experiments are conducted in a Google Colab environment using a Python 3-based infrastructure. Given the dataset size of 25.88 GB, memory mapping (mmap) is employed to avoid exceeding Google Compute Engine RAM limitations. This approach allows the data to be accessed directly from disk in batches during training without loading the entire dataset into memory, thereby enabling efficient lazy loading and optimized memory management.
3.3. Model Architecture and Training Parameters
A multimodal hybrid architecture is designed for the classification task. The model consists of two main branches:
IQ Branch: This branch takes raw IQ signal vectors of shape [B, 2, 16,384] as input. It progressively reduces the signal to [B, 256, 64] through a stem composed of four stride-2 convolutional layers. Subsequently, three residual blocks (ResBlock1D) are applied for deep feature extraction, and an Adaptive Average Pooling (AdaptiveAvgPool1d) layer produces a 256-dimensional embedding vector of shape [B, 256].
Spectrogram Branch: This branch processes spectrogram images of shape [B, 2, 128, 128] using Depthwise Separable Convolution blocks inspired by the EfficientNet architecture. Through five stages of stride-2 downsampling, the spatial resolution is reduced from 128 × 128 to 4 × 4, and an Adaptive Average Pooling layer generates a 256-dimensional embedding vector of shape [B, 256].
The embeddings obtained from both branches are concatenated to form a 512-dimensional fusion vector [B, 512]. This vector is then passed through two fully connected layers with Layer Normalization, GELU activation, and Dropout (p = 0.4), followed by a final classification layer with seven output classes. Convolutional layers are initialized using Kaiming Normal initialization, while fully connected layers are initialized using Xavier Uniform initialization.
The model is implemented using PyTorch version 2.11.0 (CPU). Training is conducted with a batch size of 64 over 40 epochs. To mitigate gradient explosion, gradient clipping (max norm = 1.0) is applied. To reduce overfitting, label smoothing (α = 0.05) is incorporated into the CrossEntropyLoss function, and the AdamW optimizer is used. The learning rate is gradually decreased over 40 epochs using a cosine annealing schedule (CosineAnnealingLR, η_min = 1 × 10−6). Mixed precision training (AMP) is employed during training to improve computational efficiency.
4. Results
In this study, the data obtained through the implementation of a multimodal algorithmic framework designed for the detection of UAV presence and model identification via RF signals were analyzed using the metrics of Loss, Accuracy, Recall, and F1-Score. The experiments were conducted across a total of seven classes, comprising six distinct drone models and one noise class.
The loss values reported during training and validation were calculated using the CrossEntropyLoss function with label smoothing, where the smoothing factor was set to 0.05. For each epoch, the loss value of each mini-batch was weighted by the corresponding batch size and then divided by the total number of samples, thereby obtaining an average loss over all samples. Accuracy was calculated as the ratio of correctly classified samples to the total number of evaluated samples. The predicted class for each sample was determined by selecting the class with the highest output logit value. In the final test evaluation, overall accuracy was computed using the predicted and ground-truth labels, while precision, recall, and F1-score were obtained on a class-wise basis and summarized using weighted average values.
4.1. Model Performance and Classification Accuracy
The proposed hybrid architecture demonstrated a stable and consistent training process. Containing 3,845,671 trainable parameters, the model achieved an overall accuracy of 87.42% on the test dataset after 40 training epochs. The parameter distribution across model components is summarized in
Table 1.
The variations in loss and accuracy throughout the training process are illustrated in
Figure 1. The adoption of the AdamW optimization algorithm with CosineAnnealingLR scheduler and Gradient Clipping techniques enabled stable convergence and high accuracy levels even under noisy data conditions.
Detailed classification performance metrics are presented in
Table 2.
The confusion matrix, which provides detailed insight into inter-class discrimination performance and error distribution, is presented in
Figure 2. The results indicate that the model exhibits strong discriminative capability despite the spectral similarities among different drone models.
The class-wise recall scores derived from the confusion matrix are further visualized in
Figure 3. The results indicate that all classes achieved a recall of 0.87, with Drone_1 exhibiting the lowest precision (0.79) due to its higher spectral similarity to other drone classes.
4.2. SNR-Based Performance Analysis
The SNR-based accuracy curve generated within the scope of this study (
Figure 4) demonstrates the noise robustness of the proposed model. To numerically support the graphical results presented in
Figure 4, a detailed performance breakdown is provided in
Table 3.
The analyses revealed that, within negative-SNR ranges (between −20 dB and 0 dB), classification accuracy varied from 64.52% to 82.40% due to the dominance of noise. In particular, the lowest SNR interval, −20 dB to −15 dB, produced the lowest accuracy value of 64.52%, indicating that the model faces greater difficulty when the noise power is considerably higher than the useful UAV signal. However, the gradual increase in accuracy from 69.80% in the −15 dB to −10 dB range to 82.40% in the −5 dB to 0 dB range shows that the proposed multimodal architecture progressively recovers discriminative signal patterns as the signal quality improves. This trend suggests that combining raw in-phase and quadrature signal features with spectrogram-based time–frequency representations provides complementary information under noisy conditions.
A clear performance transition is observed after 0 dB, where the classification accuracy exceeds 89%. This indicates that the model becomes more reliable when the useful signal becomes comparable to or stronger than the noise level. The accuracy further increases to 94.05% in the +5 dB to +10 dB range and reaches 99.10% in the +15 dB to +20 dB range. At higher SNR intervals, namely +20 dB to +25 dB and +25 dB to +30 dB, the model achieves 100% accuracy, demonstrating that the remaining classification errors mainly originate from highly noisy low-SNR conditions rather than from an inherent inability of the model to distinguish between UAV classes. Overall, these results confirm that the proposed multimodal CNN-based fusion strategy improves robustness across a wide SNR range and is particularly valuable for UAV detection scenarios where RF signal quality may fluctuate significantly.
5. Conclusions
This study demonstrated that RF signals emitted by UAVs can be autonomously classified with high accuracy through a multimodal hybrid DL architecture that jointly processes both raw IQ data and spectrogram representations. Experimental results showed that the proposed model, developed across seven different classes (six drone models and one noise class), achieved an accuracy of 87.42% and an F1-score of 0.87, while maintaining consistent detection performance within noisy frequency bands. This RF-based approach provides a passive, cost-effective, and reliable solution for the detection and classification of micro/mini UAVs that are often difficult to identify using conventional radar and optical systems, particularly in noisy and interference-heavy RF environments.
For future work, the incorporation of data augmentation strategies and transfer learning approaches is recommended, along with ablation studies to quantify the individual contributions of the in-phase and quadrature signal branch and the spectrogram branch. In addition, benchmarking inference latency on edge-oriented hardware would provide further insight into the real-time deployment potential of the proposed model. Evaluations on larger-scale datasets involving a wider variety of UAV brands and models are also encouraged, particularly under open-set recognition scenarios where previously unseen drone models may be encountered. Furthermore, conducting real-time field experiments using Software-Defined Radio platforms would enhance the practical applicability of the proposed methodology and provide a more comprehensive assessment of its operational performance in real-world environments. Finally, explainable artificial intelligence techniques may be incorporated to visualize model decision-making and improve the interpretability of the proposed detection framework.
Author Contributions
Conceptualization, T.T., S.G., A.K. and I.B.; methodology, T.T. and S.G.; software, S.G. and A.K.; validation, T.T., S.G., A.K. and I.B.; formal analysis, T.T. and A.K.; investigation, S.G. and A.K.; resources, A.K. and I.B.; data curation, T.T. and S.G.; writing—original draft preparation, T.T. and A.K.; writing—review and editing, T.T., S.G., A.K. and I.B.; visualization, S.G. and A.K.; supervision, I.B. and S.G.; project administration, I.B. and T.T.; funding acquisition, I.B. All authors have read and agreed to the published version of the manuscript.
Funding
This work is financed by the European Union-NextGenerationEU, through the National Recovery and Resilience Plan of the Republic of Bulgaria, project № BG-RRP-2.013-0001.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
The dataset used in this study is publicly available from the original source reported in Glüge et al. [
6], accessed on 3 February 2026.
Conflicts of Interest
The authors declare no conflicts of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| DL | Deep Learning |
| CNN | Convolutional Neural Network |
| UAV | Unmanned Aerial Vehicles |
| NLOS | Non-Line-of-Sight |
| RF | Radio Frequency |
| SNR | Signal-to-Noise Ratio |
References
- Rojhani, N.; Shaker, G. Comprehensive Review: Effectiveness of MIMO and Beamforming Technologies in Detecting Low RCS UAVs. Remote Sens. 2024, 16, 1016. [Google Scholar] [CrossRef] [Scilit]
- Ezuma, M.; Anjinappa, C.K.; Funderburk, M.; Guvenc, I. Radar Cross Section Based Statistical Recognition of UAVs at Microwave Frequencies. IEEE Trans. Aerosp. Electron. Syst. 2022, 58, 27–46. [Google Scholar] [CrossRef] [Scilit]
- Medaiyese, O.O.; Ezuma, M.; Lauf, A.P.; Adeniran, A.A. Hierarchical Learning Framework for UAV Detection and Identification. IEEE J. Radio Freq. Identif. 2022, 6, 176–188. [Google Scholar] [CrossRef] [Scilit]
- Nemer, I.; Sheltami, T.; Ahmad, I.; Yasar, A.U.-H.; Abdeen, M.A.R. RF-Based UAV Detection and Identification Using Hierarchical Learning Approach. Sensors 2021, 21, 1947. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yan, X.; Fu, T.; Lin, H.; Xuan, F.; Huang, Y.; Cao, Y.; Hu, H.; Liu, P. UAV Detection and Tracking in Urban Environments Using Passive Sensors: A Survey. Appl. Sci. 2023, 13, 11320. [Google Scholar] [CrossRef] [Scilit]
- Glüge, S.; Nyfeler, M.; Ramagnano, N.; Horn, C.; Schüpbach, C. Robust Drone Detection and Classification from Radio Frequency Signals Using Convolutional Neural Networks. In Proceedings of the 15th International Joint Conference on Computational Intelligence, Rome, Italy, 13–15 November 2023; pp. 496–504. [Google Scholar] [CrossRef] [Scilit]
- Alam, S.S.; Chakma, A.; Rahman, M.H.; Bin Mofidul, R.; Alam, M.M.; Utama, I.B.K.Y.; Jang, Y.M. RF-Enabled Deep-Learning-Assisted Drone Detection and Identification: An End-to-End Approach. Sensors 2023, 23, 4202. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Allahham, M.S.; Khattab, T.; Mohamed, A. Deep Learning for RF-Based Drone Detection and Identification: A Multi-Channel 1-D Convolutional Neural Networks Approach. In Proceedings of the 2020 IEEE International Conference on Informatics, IoT, and Enabling Technologies (ICIoT), Doha, Qatar, 2–5 February 2020; pp. 112–117. [Google Scholar] [CrossRef] [Scilit]
- Al-Sa’d, M.F.; Al-Ali, A.; Mohamed, A.; Khattab, T.; Erbad, A. RF-based drone detection and identification using deep learning approaches: An initiative towards a large open source drone database. Futur. Gener. Comput. Syst. 2019, 100, 86–97. [Google Scholar] [CrossRef] [Scilit]
- Ezuma, M.; Erden, F.; Kumar Anjinappa, C.; Ozdemir, O.; Guvenc, I. Detection and Classification of UAVs Using RF Fingerprints in the Presence of Wi-Fi and Bluetooth Interference. IEEE Open J. Commun. Soc. 2020, 1, 60–76. [Google Scholar] [CrossRef] [Scilit]
- Yang, S.; Luo, Y.; Miao, W.; Ge, C.; Sun, W.; Luo, C. RF Signal-Based UAV Detection and Mode Classification: A Joint Feature Engineering Generator and Multi-Channel Deep Neural Network Approach. Entropy 2021, 23, 1678. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhao, T.; Domae, B.W.; Steigerwald, C.; Paradis, L.B.; Chabuk, T.; Cabric, D. Drone RF Signal Detection and Fingerprinting: UAVSig Dataset and Deep Learning Approach. In Proceedings of the MILCOM 2024—2024 IEEE Military Communications Conference (MILCOM), Washington, DC, USA, 28 October–1 November 2024; pp. 431–436. [Google Scholar] [CrossRef] [Scilit]
- Zhang, H.; Li, T.; Li, Y.; Li, J.; Dobre, O.A.; Wen, Z. RF-Based Drone Classification Under Complex Electromagnetic Environments Using Deep Learning. IEEE Sens. J. 2023, 23, 6099–6108. [Google Scholar] [CrossRef] [Scilit]
- Elyousseph, H.; Altamimi, M. Robustness of Deep-Learning-Based RF UAV Detectors. Sensors 2024, 24, 7339. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Frid, A.; Ben-Shimol, Y.; Manor, E.; Greenberg, S. Drones Detection Using a Fusion of RF and Acoustic Features and Deep Neural Networks. Sensors 2024, 24, 2427. [Google Scholar] [CrossRef] [Scilit] [PubMed]
| Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |