Next Article in Journal
CCBA: Dynamic Scheduling Algorithm for Jammer Resources in Strong Electromagnetic Interference Environment
Previous Article in Journal
A Multidimensional Maturity Model for the Metaverse: Stages, Dimensions and Architectural Alignment
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Design of Network Traffic Analysis Models Based on Deep Neural Networks

College of Software Engineering, Zhengzhou University of Light Industry, Zhengzhou 450002, China
*
Author to whom correspondence should be addressed.
Future Internet 2026, 18(3), 152; https://doi.org/10.3390/fi18030152
Submission received: 19 January 2026 / Revised: 9 March 2026 / Accepted: 12 March 2026 / Published: 16 March 2026
(This article belongs to the Section Cybersecurity)

Abstract

The proliferation of next-generation Internet infrastructures and the Internet of Things (IoT) has exponentially increased network traffic complexity. While deep learning (DL)-based intrusion detection systems (IDSs) show immense potential, they persistently suffer from challenges including high computational overhead, vanishing gradients in deep architectures, and acute sensitivity to noise. Consequently, these issues impede their real-time deployment in resource-constrained edge computing environments. To overcome these limitations, we propose a novel, lightweight, and robust intrusion detection framework based on deep neural networks (DNNs). Initially, we employ a Robust Scaler-based statistical preprocessing strategy to supersede traditional Z-score standardization, effectively mitigating the adverse impacts of outliers and burst traffic noise. Subsequently, we design an advanced architecture that integrates self-normalizing residual blocks with a channel attention mechanism. Leveraging compressed hidden layers alongside the Scaled Exponential Linear Unit (SELU) activation function, this architecture not only mitigates the vanishing gradient problem but also amplifies critical traffic features. Concurrently, it achieves a substantial reduction in both parameter count and inference latency. Furthermore, we introduce a cosine annealing strategy to dynamically adjust the learning rate during training, thereby facilitating the model’s escape from local optima and accelerating convergence. Extensive experiments on standard benchmark datasets demonstrate that our proposed framework achieves superior detection accuracy while maintaining exceptional computational efficiency compared to state-of-the-art baselines.

1. Introduction

1.1. Purpose

With the rapid development of Internet technology, network scales are continuously expanding, accompanied by explosive growth in network traffic. As the Internet processes vast amounts of private user data, this information is highly vulnerable to diverse attacks from both internal and external intruders [1]. Consequently, network traffic analysis plays a pivotal role in critical domains such as network management, performance optimization, and cybersecurity assurance.
While advanced architectures such as convolutional neural networks (CNNs) and long short-term memory (LSTM) networks have shown remarkable performance in raw payload analysis, they inherently introduce massive computational overhead. For pre-aggregated tabular flow statistics, their spatial and sequential inductive biases are structurally redundant. Therefore, there is an urgent need for an architecture that naturally fits tabular data while avoiding the extreme parameter inflation of CNNs and LSTMs.
This study aims to conduct an in-depth exploration of network traffic analysis models based on deep neural networks (DNNs) [2]. By meticulously designing, training, and optimizing the DNN architecture, this research enables the comprehensive extraction of critical features from network traffic data, thereby achieving precise classification and effective prediction of network traffic patterns [3]. Specifically, this research is dedicated to addressing the following fundamental issues: how to design an appropriate DNN architecture tailored to the specific characteristics and analytical requirements of network traffic data; how to select and preprocess network traffic data to enhance data quality, usability, and observability, providing a solid data foundation for model training; how to optimize the DNN training process to accelerate convergence and improve generalization capabilities while mitigating issues such as overfitting [3] and underfitting [4]; and how to comprehensively evaluate and validate the performance of the proposed DNN-based network traffic analysis model to ensure its effectiveness and reliability in real-world applications.

1.2. Methods and Innovations

To ensure the scientific rigor and validity of this study, a comprehensive methodology is employed, detailed as follows:
Data Collection and Preprocessing: The proposed model utilizes the CICIDS2017 [5] and CIC-IoT-2023 [6] datasets. These are widely recognized benchmark datasets employed to evaluate and compare the detection accuracy and training efficiency of the models.
The primary innovations and contributions of this study regarding the DNN-based network traffic analysis model are summarized as follows:
  • Self-Normalizing Residual Architecture: Building upon the foundational DNN architecture, this study introduces a self-normalizing residual block based on the SELU activation function [7,8]. This design effectively resolves the vanishing gradient problem inherent in deep networks. Concurrently, by restricting each hidden layer to merely 20 neurons, the model maintains a remarkably low parameter count, thereby satisfying the stringent requirements for real-time detection in edge devices or high-speed networks.
  • Channel Attention Mechanism: To further enhance the model’s capacity to extract high-dimensional network traffic features, a channel attention (squeeze-and-excitation) mechanism is incorporated. Originally proposed by Hu et al. [9], this mechanism adaptively recalibrates feature responses by explicitly modeling the interdependencies between feature channels.
  • Robust Statistical Preprocessing: Addressing the prevalent issues of outliers and burst traffic in network data—where traditional Z-score standardization is severely compromised by extreme values—this paper adopts a Robust Scaler strategy based on the median and the interquartile range (IQR).
  • Cosine Annealing Optimization: During the model training phase, a cosine annealing learning rate scheduling strategy is implemented. Initially introduced by Loshchilov et al. [10], recent studies [11] have demonstrated that integrating the cosine annealing mechanism in complex IoT intrusion detection tasks can significantly accelerate model convergence. It achieves higher F1-scores and accuracy with minimal hyperparameter tuning costs [12], which strongly aligns with the performance enhancements observed in the ablation studies in this work.
While convolutional neural networks (CNNs) and long short-term memory (LSTM) networks have been widely employed in hybrid intrusion detection systems, they are structurally suboptimal for flow-based statistical traffic analysis. The selection of a deep neural network (DNN) architecture in this study is motivated by the specific inductive biases of the models and the nature of the input data:
  • CNNs rely on strong spatial inductive biases, making them ideal for grid-like data or raw payload byte sequences. However, the data utilized in this study consist of tabular, flow-level statistical features (e.g., flow duration, forward/backward packet lengths). These features are inherently heterogeneous and lack spatial correlation. Forcing such tabular data into pseudo-images for convolutional filtering disrupts feature independence and introduces unnecessary computational overhead.
  • LSTMs are designed to capture temporal dependencies in sequential data. Although raw network packets arrive sequentially, our framework processes pre-aggregated flow statistics. The temporal dynamics of the network flows have already been mathematically encoded into features such as inter-arrival times and active/idle durations. Consequently, applying recurrent architectures to these aggregated, static feature vectors is highly redundant and significantly inflates inference latency, which is detrimental to real-time deployment.
  • Therefore, a DNN architecture is naturally the most suitable choice for processing tabular statistical data. To overcome the inherent limitations of traditional DNNs, such as vanishing gradients and feature neglect as network depth increases, our proposed framework introduces self-normalizing residual blocks and a channel attention mechanism. This tailored DNN architecture efficiently recalibrates the importance of various statistical features, achieving a superior balance between detection accuracy and computational efficiency compared to heavy CNN or LSTM-based hybrid models.

2. Construction of the DNN-Based Network Traffic Analysis Model

2.1. Data Sources and Preprocessing

This study primarily utilized the CICIDS2017 and CIC-IoT-2023 datasets. These are widely recognized and publicly available datasets within the cybersecurity domain, characterized by extensive network traffic data that encompasses a diverse array of network application scenarios and attack types. They provide comprehensive data support for the training and validation of the proposed model [6]. Furthermore, CIC-IoT-2023 is a recent dataset specifically tailored for Internet of Things (IoT) environments. Evaluating the proposed model on this dataset effectively demonstrates that its lightweight design and real-time processing capabilities highly align with the stringent operational requirements of IoT devices.

2.1.1. Feature Selection Based on Gini Importance and Data Partitioning

To facilitate lightweight deployment on resource-constrained IoT edge nodes, reducing the input dimensionality is an indispensable prerequisite. Utilizing all available flow features—which typically encompass more than 40 to 70 dimensions depending on the dataset—would result in an excessive number of parameters in the initial dense layer of the neural network, thereby significantly increasing computational overhead and inference latency.
Consequently, this study employs a random forest (RF) classifier to rigorously evaluate feature importance [13]. The RF algorithm calculates the Gini Importance, also known as the Mean Decrease in Impurity, for each individual feature [14]. The features are subsequently ranked in descending order based on their calculated importance scores, and the top 15 features are selected to form the optimal feature subset. This critical step effectively eliminates redundant and context-specific attributes (e.g., timestamps and IP identifiers) that do not contribute to generalized intrusion detection, thus substantially reducing the computational burden on the subsequent DNN model.
Table 1 presents the top 15 selected features obtained through this RF-based importance screening process.
To ensure a comprehensive evaluation of the model’s performance across different data distributions, the preprocessed and feature-extracted network traffic data are partitioned into training, validation, and testing sets. In this study, the dataset is split utilizing an 8:2 ratio; specifically, 80% of the data is allocated for training the model, while the remaining 20% is reserved for testing. This specific partitioning ratio was selected based on a comprehensive consideration of optimizing the model’s training efficacy while ensuring the maximal utilization of the available data.
To further validate the model’s performance and stability during the training process, 10% of the training set is allocated as a validation set. During learning rate tuning, various learning rate values are systematically applied. The model is trained on the training set and subsequently evaluated on the validation set. The learning rate that yields the optimal performance on the validation set is then selected as the final hyperparameter.

2.1.2. Normalization

Network traffic datasets, particularly those encompassing Denial of Service (DoS/DDoS) attacks, frequently manifest extreme outliers in their statistical features. Conventional normalization techniques, such as Z-score (standard scaling) or min–max scaling, are highly susceptible to these outliers. This susceptibility induces a distortion within the feature space, subsequently compressing the distribution space of normal traffic patterns.
To bolster the model’s robustness against noise and burst traffic, the Robust Scaler from the Scikit-learn library [15] was adopted to standardize the retained features. In contrast to the Z-score approach, which is heavily dependent on the mean and standard deviation, the Robust Scaler leverages the median and interquartile range (IQR). The scaling formula for a given feature vector x is defined as follows:
x ′ = x − Q 2 ( x ) Q 3 ( x ) − Q 1 ( x )
where Q 1 ( x ) ,   Q 2 ( x ) and Q 3 ( x ) denote the 25th percentile, the median (50th percentile), and the 75th percentile of the feature, respectively. This approach ensures that extreme outliers do not dominate the scaling process, thereby effectively preserving the inherent distributional characteristics of the legitimate traffic data.

2.2. Model Architecture Design

In accordance with the specific characteristics and requirements of network traffic analysis tasks, the proposed DNN-based model is structured into three primary components: the input and preprocessing layer, the stacked residual-attention blocks, and the output classification layer.
Initially, feature selection is performed on the dataset. To align with the lightweight constraint of utilizing only 20 neurons in the hidden layers, a random forest algorithm is employed to extract the top 15 most discriminative features, thereby effectively reducing the input dimensionality. Following this, robust statistical scaling methods are applied to clean and normalize the raw traffic data, mitigating the adverse effects of extreme outliers.
The preprocessed features are subsequently fed into a deep neural network featuring a bottleneck architecture. This network seamlessly integrates residual connections and attention mechanisms to accurately classify the network traffic as either benign or malicious. The comprehensive architecture of the proposed model is illustrated in Figure 1.

2.2.1. Activation Function

Standard deep neural networks (DNNs) typically employ wide hidden layers. In contrast, the proposed model restricts the hidden layers to merely 20 neurons to meet lightweight constraints. To ensure the training stability of such a narrow network architecture, the Scaled Exponential Linear Unit (SELU) [16] activation function is adopted. The SELU possesses a self-normalizing property that inherently maintains a mean of zero and variance of one across layers, thereby eliminating the necessity for explicit batch normalization (BN) layers. This characteristic substantially reduces both the parameter count and the computational overhead during inference, making the model highly suitable for deployment on resource-constrained IoT edge devices. Furthermore, given the limited number of neurons in the proposed architecture, traditional activation functions like ReLU are highly susceptible to the “dying ReLU” problem. This phenomenon can lead to a severe reduction in effective feature dimensionality. Consequently, the adoption of SELU mitigates this risk and significantly accelerates the convergence rate of the feature values. The mathematical expression for the SELU function is defined as follows:
f ( x ) = λ { x x > 0 α ( e x − 1 ) x ≤ 0
where λ ≈ 1.0507 and α ≈ 1.6733 are predefined scaling constants.

2.2.2. Residual-Attention Block

Within the compressed 20-dimensional feature space, increasing the network depth intrinsically elevates the risk of information degradation and the vanishing gradient problem. To mitigate these challenges, a residual-attention block is introduced and stacked multiple times throughout the network architecture. The residual-attention block is displayed in Figure 2.
Channel Attention Mechanism [7]: Inspired by the squeeze-and-excitation (SE) network, a lightweight channel attention module is integrated into the architecture. This module explicitly models the interdependencies across the 20 feature channels. It adaptively assigns higher weights to channels containing critical attack features while actively suppressing noisy channels. The attention weight vector, denoted as s, is computed through consecutive squeeze (dimensionality reduction) and excitation (dimensionality expansion) operations, defined as follows:
s = σ ( W u p δ ( W d o w n x ) )
x o u t = x ⨂ s
where W d o w n and W u p represent the weights of the dimensionality reduction and expansion layers, respectively. δ denotes the ReLU activation function, and σ represents the Sigmoid function. The final output x o u t is obtained by scaling the input feature map x with the attention weight vector s using element-wise multiplication (⨂).
Residual Connection [8]: To facilitate gradient flow, a residual connection (identity mapping) is introduced in parallel with the attention transformation. The final output y of the block is computed as follows:
y = x o u t + x
This architectural design ensures that gradients can propagate directly to earlier layers, thereby enabling the effective training of deep and narrow networks.

2.2.3. Optimization Strategy

Optimization algorithms [17] are responsible for iteratively adjusting the model parameters during the training process of deep neural networks. Their primary objective is to continuously minimize the empirical loss on the training dataset, thereby enhancing the model’s accuracy and generalization capabilities. To address the prevalent issues of slow convergence and the tendency to fall into local minima, the proposed framework integrates the Adam (Adaptive Moment Estimation) optimizer with a Cosine Annealing Warm Restarts [12] learning rate scheduling strategy.
Initially, Adam is employed as the foundational optimizer. Given the sparse and non-stationary nature of IoT traffic data, Adam designs adaptive learning rates for different parameters by computing the first-order moment estimate (the mean of the gradients) and the second-order moment estimate (the uncentered variance of the gradients). The mathematical update rules for the parameter θ at time step t are defined as follows:
g t = ∇ θ   J ( θ t − 1 )
m t = β 1 m t − 1 + ( 1 − β 1 ) g t
v t = β 2 v t − 1 + ( 1 − β 2 ) g t 2
To counteract the initialization bias towards zero during the early stages of training, bias-corrected moment estimates are computed:
m t ^ = m t 1 − β 1 t
v t ^ = v t 1 − β 2 t
Finally, the parameter θ is updated using the corrected estimates:
θ t = θ t − 1 −   η t v t ^ + ε m t ^
Here, m t ^ and v t ^ denote the bias-corrected first and second moment estimates, respectively, and ε is a smoothing term added to ensure numerical stability. Although Adam can adaptively adjust the update direction, its global learning rate η t is typically fixed or monotonically decaying, which constrains the model’s capability to escape saddle points.
To overcome the limitations of a fixed learning rate, we integrate the Cosine Annealing Warm Restarts strategy into the η t term of the Adam update formula. Instead of employing a monotonic decay, this strategy allows the learning rate to vary periodically following a cosine function. The learning rate η t at epoch t is computed as follows:
η t = η m i n + 1 2 ( η m a x − η m i n ) ( 1 + cos ( T c u r T m a x ) π )
where η m a x and η m i n represent the maximum and minimum learning rates, respectively, and T m a x denotes the length of the annealing cycle. This strategy periodically resets the learning rate, enabling the optimizer to escape sharp local minima and converge towards flatter solutions.

3. Experiments and Results

3.1. Selection of Evaluation Indicators

To comprehensively and accurately evaluate the performance of the proposed DNN-based network traffic analysis model, this paper adopts a standard evaluation framework based on the confusion matrix [18]. This framework encompasses several critical evaluation metrics, including accuracy, recall, and the F1-score, which play a vital role in assessing the efficacy of traffic analysis models. The specific definitions and calculation formulas for these metrics are detailed in Table 2.
TP (True Positive) is the number of samples that were actually positive and were correctly predicted to be positive by the model.
TN (True Negative) is the number of examples that are actually negative and are correctly predicted as negative by the model.
FP (False Positive) stands for false positives. This is the number of examples that were actually negative but were incorrectly predicted to be positive by the model.
FN (False Negative) stands for false negative, which is the number of examples that were actually positive but were incorrectly predicted to be negative by the model.
Accuracy refers to the proportion of the number of samples correctly predicted by the model to the total number of samples, and it is one of the most intuitive performance indicators, reflecting the accuracy of the model in classifying the overall sample. It is calculated as follows:
A c c u r a c y = T P + T N T P + T N + F P + F N
Recall, also known as the true-positive rate (TPR) or sensitivity, is the proportion of examples that are actually positive and are correctly predicted as positive by the model; it reflects the ability of the model to capture positive examples. In network traffic analysis, recall is particularly important for identifying abnormal traffic because missing abnormal traffic data may lead to serious security problems. If the model correctly identifies 80 out of 100 abnormal traffic samples, the recall is 80 ÷ 100 × 100% = 80%. It is calculated as follows:
R e c a l l = T P T P + F N
The F1-score is the harmonic mean of precision and recall, which takes into account the balance between precision and recall. Precision refers to the proportion of samples predicted as the positive class that are actually positive, and it measures the accuracy of the prediction result, which is calculated by the following formula:
P r e c i s i o n = T P T P + F P
The F1 value ranges from 0 to 1, with 1 indicating perfect precision and recall, and a higher F1 value indicates better overall performance of the model. For example, when the precision is 0.8 and the recall is 0.7, the F1 value is approximately 0.74. It is calculated as follows:
F 1 = 2 × P r e c i s i o n × R e c a l l P r e c i s i o n + R e c a l l
Here, we introduce the false-positive rate (FPR) to measure the fraction of normal connection records that are incorrectly labeled as attacks. A low false-positive rate indicates that the model has few false positives and is defined as follows:
F P R = F P F P + F N
In network traffic analysis, these evaluation metrics complement each other and can comprehensively evaluate the performance of the model from different perspectives. Through the comprehensive use of these evaluation indicators, the effectiveness and reliability of DNN-based network traffic analysis models in practical applications can be more accurately judged.

3.2. Experimental Setup and Baseline Models

To validate the stability of the proposed model in terms of lightweight design and robustness, comparative experiments were conducted against three traditional baseline models. These baselines encompass both traditional machine learning algorithms and mainstream deep learning approaches.
  • Random Forest (RF) [13]: Representing traditional machine learning algorithms, RF is widely applied in the field of intrusion detection due to its strong interpretability and high accuracy. In the comparative experiments, the number of estimators (decision trees) was set to 100, and Gini impurity was utilized as the criterion for splitting nodes.
  • One-Dimensional Convolutional Neural Network (1D-CNN) [19]: This model represents the state-of-the-art deep learning methodology currently employed in traffic classification. CNNs are capable of automatically extracting local spatial features from traffic sequences through convolution kernels. For the comparative experiments, the 1D-CNN architecture comprises two convolutional layers (with 32 and 64 filters, respectively), coupled with a max-pooling layer and a fully connected layer. The objective of this comparison is to ensure that the proposed DNN can achieve equivalent or superior detection accuracy compared to the 1D-CNN in the training results.
  • Baseline DNN [20]: This model possesses a depth similar to the proposed novel DNN (4 layers) but features a wider network architecture, with each hidden layer containing 64 neurons. It utilizes the conventional ReLU [16] activation function alongside batch normalization, and it distinctly lacks both residual connections and attention mechanisms. This standard deep fully connected network serves as a baseline to explicitly validate the effectiveness of the architectural improvements introduced in our proposed model.
The experimental environment was constructed using the TensorFlow 2.6 framework. The hardware platform was configured with an Intel Core i7 CPU and 16 GB of RAM, and notably, GPU acceleration was disabled during the evaluation. This setup strictly simulates resource-constrained environments. The evaluation metrics encompass accuracy, precision, recall, and the F1-score. Crucially, inference time and parameter count were also evaluated, as these metrics are paramount for real-world deployment in IoT scenarios.

3.3. Experimental Data Analysis Performance Comparison Analysis

To comprehensively evaluate the performance of the proposed model, extensive tests were conducted on the CIC-IoT-2023 dataset. Table 3 presents the detailed performance metrics of the various models evaluated in the multi-class classification task. Furthermore, Table 4 illustrates the comparative performance data obtained when evaluating the models using the CICIDS2017 dataset.
Based on the empirical data presented in the two tables above, the following analytical conclusions can be drawn:
  • Superiority over Baseline DNN: The proposed model consistently outperforms the Baseline DNN across all four primary performance metrics: accuracy, precision, recall, and the F1-score. Notably, the detection accuracy exhibits a significant improvement of approximately 28.01%. This enhancement demonstrates that a simple stacking of fully connected layers, as seen in the Baseline DNN, is highly susceptible to information loss when processing high-dimensional traffic features. In contrast, the proposed model effectively mitigates the vanishing gradient problem by incorporating residual connections and a channel attention mechanism. These architectural innovations enhance the model’s ability to capture critical attack features while simultaneously demonstrating that the refined architecture significantly improves overall modeling efficiency.
  • Comparison with 1D-CNN: When compared to the 1D-CNN, the proposed model achieves a detection accuracy that is essentially on par with the Baseline CNN, showing no distinct disadvantage in precision. However, a critical advantage emerges when analyzing the model complexity: the parameter count of the proposed DNN is only approximately 5% of that required by the 1D-CNN. Achieving comparable performance with such a drastically reduced parameter footprint confirms that the “Narrow Bottleneck + Attention” design is exceptionally efficient for resource-constrained traffic analysis tasks.
  • Comparison with Random Forest (RF): While the random forest algorithm achieves the highest overall accuracy, its practical utility in IoT environments is severely limited by its substantial model size, as detailed in the subsequent analysis of Table 5. The proposed DNN maintains a high level of accuracy that is only marginally lower than RF, yet it boasts an extremely compact file size of only 12 KB. Consequently, in real-world IoT deployment scenarios where memory and storage are at a premium, our proposed model represents a superior and more viable alternative.
For practical deployment in IoT and network environments, the proposed model must not only provide precise detection but also demonstrate superior computational efficiency. The objective is to minimize computational overhead while maintaining high detection performance, thereby achieving an optimal trade-off between accuracy and efficiency.
To evaluate the lightweight nature of the architecture, Table 5 provides a comprehensive comparison of the spatial complexity (represented by parameter count) and temporal complexity (represented by inference latency) across the evaluated models.
The parameter count for the Random Forest (RF) model is denoted as "N/A" (Not Applicable) because RF is an ensemble of decision trees and does not optimize a fixed set of trainable weights and biases like neural networks. Instead, its complexity is determined by hyperparameters such as the number of trees and their maximum depth.
The parameter count of DNN is approximately 2339, which represents a reduction of about 91% compared to 1D-CNN and 77% compared to the Baseline DNN. This efficiency is primarily attributed to our bottleneck design, which strictly limits the hidden layers to 20 neurons. Such a minimal model size (only 56.91 KB) allows it to be easily loaded into the on-chip cache (L1 Cache) of a Raspberry Pi or even a microcontroller (MCU), significantly reducing memory access latency. In contrast, while random forest (RF) offers relatively fast inference, its model files typically reach several MBs—depending on the depth and number of trees—which is often unacceptable for memory-constrained IoT sensor nodes.
DNN achieves an inference speed of 30 μs/sample, which is 15% faster than 1D-CNN.
Contribution of SELU: Although the Baseline DNN has a simple structure, it requires additional computational steps during inference due to the inclusion of batch normalization (BN) layers. In contrast, DNN utilizes the SELU activation function to achieve self-normalization, effectively eliminating the need for BN layers and further shortening the computational path.
Comparison with RF: While random forest (RF) demonstrates competitive speed on CPUs, DNN exhibits superior throughput and possesses greater potential for hardware acceleration.
Synthesizing the results from Table 3 and Table 4, it can be inferred that DNN does not blindly pursue the highest possible accuracy. Instead, it minimizes computational costs while maintaining state-of-the-art (SOTA) detection precision (>99%). This optimal trade-off between accuracy and efficiency proves that DNN is the most suitable intrusion detection system (IDS) solution for deployment on IoT edge nodes in the future Internet.
The empirical results further validate our architectural choice. As demonstrated, the proposed DNN achieves an F1-score comparable to or even better than the computationally heavy 1D-CNN baseline, but with only a fraction of its parameter count and inference latency. This confirms our theoretical rationale: for pre-aggregated tabular traffic features, a well-optimized, narrow DNN combined with attention mechanisms is structurally superior to spatially or sequentially biased models like CNNs and LSTMs, achieving the ultimate trade-off between precision and computational efficiency.

3.4. Ablation Study

To validate the effectiveness of the individual components within the DNN framework—namely, robust preprocessing, the SELU activation function, residual connections, the channel attention mechanism, and the cosine annealing strategy—we designed a progressive ablation study. To guarantee a strictly fair comparison, all model variants were trained utilizing an identical set of input features and the exact same hyperparameter configurations. Commencing with a Baseline DNN, each architectural component was incrementally incorporated. The specific configurations of all model variants are detailed in Table 6.
By systematically comparing three primary evaluation metrics—specifically, accuracy, F1-score, and inference time—across the five model variants, we quantitatively analyze the incremental performance gains achieved at each progressive stage of the architectural design. The detailed experimental results are delineated in Table 7 and graphically illustrated in Figure 3.
Based on the progressive performance gains observed in the ablation experiments, the specific impact and theoretical contribution of each architectural component are analyzed as follows:
  • Impact of Robust Preprocessing: As demonstrated by the comparison between Model A and Model B in Table 7, replacing the standard Z-score normalization with the median and IQR-based Robust Scaler yields a 0.4% improvement in the F1-score. In IoT network traffic, DDoS attacks frequently generate extreme statistical outliers. Traditional Z-score scaling is highly susceptible to this mean shifting, which consequently compresses the majority of benign traffic features into an exceedingly narrow interval. Conversely, the Robust Scaler effectively preserves the underlying data distribution, enabling the model to delineate clearer classification boundaries.
  • Impact of the SELU Activation Function: The results from Model C indicate that within narrow hidden layers comprising merely 20 neurons, SELU significantly outperforms traditional ReLU. Given the “bottleneck” design of our network architecture, ReLU is prone to the “dying ReLU” (dead neurons) problem in negative regions, leading to the irreversible loss of already scarce feature information. The self-normalizing property of SELU not only prevents vanishing gradients but also ensures that all neurons remain active. Furthermore, the utilization of SELU eliminates the necessity for batch normalization (BN) layers. This maintains high detection precision while exerting virtually no negative impact on inference latency (and implicitly reduces memory access overhead).
  • Impact of Channel Attention Mechanism: Model E (the finalized DNN) achieves optimal performance through the integration of the channel attention mechanism. Although incorporating this attention computation module incurs a marginal inference time overhead of 1.33 μs, it delivers a substantial 0.7% enhancement in the F1-score. This substantiates that within a strictly constrained feature space (20 dimensions), explicitly weighting (recalibrating) feature channels is exceptionally crucial. It empowers the model to adaptively focus on the key fingerprint features of malicious attack traffic while suppressing irrelevant noise.
  • Impact of Cosine Annealing Optimization: Compared to a fixed learning rate (Fixed LR) strategy, Model E, which employs Cosine Annealing Warm Restarts, demonstrates a significantly faster convergence rate. During the initial training phases, the loss function descends much more rapidly. In the later stages, the periodic learning rate resets successfully, enabling the optimizer to escape sharp local optima, ultimately converging to a flatter and lower overall loss level. This confirms that the DNN framework is not only more accurate in anomaly detection but also highly efficient in its training process.

3.5. Performance Under Adversarial Evasion Attacks

To empirically validate the practical security contribution of the proposed framework, particularly its robustness against statistical evasion attacks, we conducted an adversarial noise injection experiment. In advanced persistent threats (APTs), attackers often inject extreme outlier packets (burst noise) to distort flow-level statistics and evade IDS detection.
We simulated this evasion tactic by randomly injecting extreme multiplier noise (scaling specific statistical features like packet length variance and inter-arrival time by a factor of 50) into 5%, 10%, and 20% of the malicious test samples. We compared the performance degradation of a standard DNN model (using standard Z-score normalization) against our proposed DNN framework (utilizing Robust Scaler based on median and IQR).
The specific values are shown in Table 8.
The experimental results reveal critical insights into model security:
  • The Vulnerability of Traditional Machine Learning (RF): While the random forest (RF) model achieves a deceptively high initial F1-score of 0.9903 on clean data, it demonstrates severe vulnerability to adversarial noise. Its performance precipitously drops to 0.6289 under 20% evasion noise. This indicates that the 99% accuracy is largely a result of overfitting to clean datasets; its hard decision boundaries are extremely brittle and can be easily bypassed by attackers slightly modifying traffic patterns.
  • The Fragility of Standard Deep Learning (CNN & Baseline DNN): Conventional deep learning models also exhibit significant security flaws. The Baseline CNN drops from 0.8954 to 0.5514, and the Baseline DNN collapses completely to 0.4135. Standard architectures lack built-in mechanisms to filter out malicious perturbations, allowing noise to propagate and amplify through the network layers, ultimately destroying the classification logic.
  • The Robustness of the Proposed DNN: In stark contrast, the proposed DNN demonstrates exceptional adversarial robustness. It starts at a realistic, non-overfitted baseline of 0.9030 and maintains a high F1-score of 0.8293 even under severe 20% evasion noise. This graceful degradation is structurally guaranteed by the proposed components: the Robust Scaler effectively neutralizes extreme adversarial outliers using median and interquartile ranges, while the network design prevents the internal amplification of adversarial perturbations.
Consequently, the proposed RRA-DNN proves to be not only highly accurate but also structurally secure, making it highly reliable for deployment in hostile, real-world network environments where traffic obfuscation and evasion attacks are frequent.

4. Conclusions

To address the critical challenges of high-dimensional feature redundancy, low recognition rates for complex attacks, and strict inference latency constraints of edge computing devices in Internet of Things (IoT) and general network environments, this paper proposes a lightweight, high-precision intrusion detection model driven by feature selection and deep architecture optimization.
Initially, this study leverages the random forest algorithm to evaluate and extract core traffic features from the CIC-IoT-2023 and CICIDS2017 datasets, effectively mitigating the risk of the curse of dimensionality. Subsequently, a Robust Scaler is introduced to eliminate the statistical distribution distortion caused by extreme outlier values inherent in malicious attack traffic. In terms of network design, this paper departs from traditional deep neural networks (DNNs), which are prone to feature space degradation. Instead, it innovatively integrates residual connections to ensure stable gradient propagation across deep layers. Concurrently, the fundamental architecture replaces the single SELU activation function with a robust Dense + BN + ReLU combination while explicitly embedding a channel attention mechanism within the deeper structures. This synergistic design achieves adaptive focusing and enhanced extraction of critical attack fingerprints.
To rigorously validate the effectiveness and superiority of the proposed framework, comprehensive ablation studies and multi-model comparative analyses were conducted on both the CIC-IoT-2023 and CICIDS2017 datasets, evaluating the model’s cross-scenario generalization capability. The experimental procedures strictly controlled random seeds and hyperparameter configurations to ensure absolute fairness. Furthermore, 1D-CNN and traditional DNN architectures were introduced as baseline models for performance benchmarking. Regarding evaluation metrics, beyond standard accuracy, special emphasis was placed on the F1-score—which more accurately reflects the true detection capability in class-imbalanced scenarios—and the single-sample inference time.
Experimental results demonstrate that the finalized optimized model (Model E) achieves a significant breakthrough in classification performance across both datasets, with its accuracy and F1-score consistently exceeding 0.995. Crucially, visual analysis via dual Y-axis line charts confirms that despite the incorporation of deep residual and attention modules, the model incurs only a marginal computational overhead (maintaining an inference time of approximately 19.76 μs). This successfully achieves an optimal trade-off between detection precision and inference efficiency in stringent IoT deployment environments.
In conclusion, the joint optimization scheme encompassing feature dimensionality reduction and network architecture proposed in this study not only establishes a highly efficient and reliable new paradigm for network security in complex IoT environments but also provides a vital theoretical foundation and practical reference for the lightweight design of future edge intelligence security systems.

Author Contributions

Conceptualization, J.C. and Y.Z.; methodology, Y.Z.; software, Y.Z.; validation, J.C. and Y.Z.; formal analysis, J.C.; investigation, Y.Z.; resources, J.C.; data curation, Y.Z.; writing—original draft preparation, Y.Z.; writing—review and editing, J.C. and Y.Z.; visualization, Y.Z.; supervision, J.C.; project administration, Y.Z.; funding acquisition, J.C. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Specialized and Creative Integration Characteristic Demonstration Courses (Second Batch) in Henan Province (Comprehensive Practice of Computer Network Technology), grant number 74; and the Key R&D and Promotion Projects in Henan Province, grant number 252102211105.

Data Availability Statement

The data presented in this study are available on request from the corresponding author.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Mukherjee, B.; Heberlein, L.T.; Levitt, K.N. Network intrusion detection. IEEE Netw. 2002, 8, 26–41. [Google Scholar] [CrossRef] [Scilit]
  2. Sze, V.; Chen, Y.H.; Yang, T.J.; Emer, J.S. Efficient processing of deep neural networks: A tutorial and survey. Proc. IEEE 2017, 105, 2295–2329. [Google Scholar] [CrossRef] [Scilit]
  3. Rimal, Y.; Sharma, N.; Alsadoon, A. The accuracy of machine learning models relies on hyperparameter tuning: Student result classification using random forest, randomized search, grid search, bayesian, genetic, and optuna algorithms. Multimed. Tools Appl. 2024, 83, 74349–74364. [Google Scholar] [CrossRef] [Scilit]
  4. Fan, Y.; Huang, H.; Han, H. Quantifying overfitting in deep learning. Anal. Appl. 2025, 23, 705–729. [Google Scholar] [CrossRef] [Scilit]
  5. Engelen, G.; Rimmer, V.; Joosen, W. Troubleshooting an intrusion detection dataset: The CICIDS2017 case study. In Proceedings of the 2021 IEEE Security and Privacy Workshops (SPW); IEEE: New York, NY, USA, 2021; pp. 7–12. [Google Scholar]
  6. Erskine, S.K. Real-time large-scale intrusion detection and prevention system (IDPS) CICIoT dataset traffic assessment based on deep learning. Appl. Syst. Innov. 2025, 8, 52. [Google Scholar] [CrossRef] [Scilit]
  7. Abdelhamid, S.; Hegazy, I.; Aref, M.; Roushdy, M. Attention-driven transfer learning model for improved IoT intrusion detection. Big Data Cogn. Comput. 2024, 8, 116. [Google Scholar] [CrossRef] [Scilit]
  8. Cui, B.; Chai, Y.; Yang, Z.; Li, K. Intrusion detection in IoT using deep residual networks with attention mechanisms. Future Internet 2024, 16, 255. [Google Scholar] [CrossRef] [Scilit]
  9. Hu, J.; Shen, L.; Sun, G. Squeeze-and-excitation networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2018; pp. 7132–7141. [Google Scholar]
  10. Loshchilov, I.; Hutter, F. Sgdr: Stochastic gradient descent with warm restarts. arXiv 2016, arXiv:1608.03983. [Google Scholar]
  11. Wang, A.; Wang, W.; Zhou, H.; Zhang, J. Network intrusion detection algorithm combined with group convolution network and snapshot ensemble. Symmetry 2021, 13, 1814. [Google Scholar] [CrossRef] [Scilit]
  12. Shao, Y.; Yang, J.; Zhou, W.; Sun, H.; Xing, L.; Zhao, Q.; Zhang, L. An improvement of Adam based on a cyclic exponential decay learning rate and gradient norm constraints. Electronics 2024, 13, 1778. [Google Scholar] [CrossRef] [Scilit]
  13. Breiman, L. Random forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef] [Scilit]
  14. Kouassi, B.M.; Ballo, A.B.; Ayikpa, K.J.; Mamadou, D.; Coulibaly, M.Z.J. Top-K Feature Selection for IoT Intrusion Detection: Contributions of XGBoost, LightGBM, and Random Forest. Future Internet 2025, 17, 529. [Google Scholar] [CrossRef] [Scilit]
  15. Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; et al. Scikit-learn: Machine learning in Python. J. Mach. Learn. Res. 2011, 12, 2825–2830. [Google Scholar]
  16. Klambauer, G.; Unterthiner, T.; Mayr, A.; Hochreiter, S. Self-normalizing neural networks. Adv. Neural Inf. Process. Syst. 2017, 30, 972–981. [Google Scholar]
  17. Abdulkadirov, R.; Lyakhov, P.; Nagornov, N. Survey of optimization algorithms in modern neural networks. Mathematics 2023, 11, 2466. [Google Scholar] [CrossRef] [Scilit]
  18. Powers, D.M. Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation. arXiv 2020, arXiv:2010.16061. [Google Scholar] [CrossRef] [Scilit]
  19. Elsayed, E.B.; Yassin, A.S.; Fahmy, H. A Novel Hybrid GWO-RFO Metaheuristic Algorithm for Optimizing 1D-CNN Hyperparameters in IoT Intrusion Detection Systems. Information 2025, 16, 1103. [Google Scholar] [CrossRef] [Scilit]
  20. Ferrag, M.A.; Maglaras, L.; Moschoyiannis, S.; Janicke, H. Deep learning for cyber security intrusion detection: Approaches, datasets, and comparative study. J. Inf. Secur. Appl. 2020, 50, 102419. [Google Scholar] [CrossRef] [Scilit]
Figure 1. DNN model structure.
Figure 1. DNN model structure.
Futureinternet 18 00152 g001
Figure 2. Residual-attention block.
Figure 2. Residual-attention block.
Futureinternet 18 00152 g002
Figure 3. Indicators of each model in the ablation experiment.
Figure 3. Indicators of each model in the ablation experiment.
Futureinternet 18 00152 g003
Table 1. The characteristics and values of the selected dataset.
Table 1. The characteristics and values of the selected dataset.
CICIDS2017ValueCIC-IOT2021Value
Target0.1054IAT0.2396
Total_Length_of_Fwd_Packets0.0584syn_count0.0561
Destination_Port0.0556Magnitue0.0549
Bwd_Packets/s0.0458psh_flag_number0.0497
Total_Fwd_Packets0.0449Protocol Type0.0475
Fwd_IAT_Mean0.0416syn_flag_number0.0406
Fwd_Packet_Length_Max0.0398Tot size0.0402
Avg Fwd Segment Size0.0383Header_Length0.0379
Average_Packet_Size0.0378fin_flag_number0.0376
Subflow_Fwd_Bytes0.0339Min0.0372
Subflow Bwd Bytes0.0309Tot sum0.0331
Subflow_Fwd_Packets0.0308AVG0.0327
Packet Length Variance0.0208ack_count0.0303
Fwd_Packet_Length_Mean0.0261rst_count0.0302
Max Packet Length0.0233MAX0.0281
Table 2. Definition of evaluation indicators.
Table 2. Definition of evaluation indicators.
Real ResultsDetection Results
TrueFalse
TrueTPFN
FalseFPTN
Table 3. Performance indicator data of each model in CIC-IOT2023.
Table 3. Performance indicator data of each model in CIC-IOT2023.
ModelAccuracyPrecisionRecallF1-Score
Random Forest99.03%98.98%99.03%98.95%
1D-CNN89.54%88.11%89.54%88.13%
Baseline DNN62.29%66.78%62.29%51.99%
DNN90.30%90.85%90.30%89.23%
Table 4. Performance indicator data of each model in CICIDS2017.
Table 4. Performance indicator data of each model in CICIDS2017.
ModelAccuracyPrecisionRecallF1-Score
Random Forest99.99%99.99%99.99%99.99%
1D-CNN99.90%99.90%99.90%99.90%
Baseline DNN98.82%98.86%98.82%98.82%
DNN99.83%99.83%99.83%99.83%
Table 5. Comparison of model complexity and size.
Table 5. Comparison of model complexity and size.
ModelParametersModel Size (KB)Latency (μs)Throughput (s/s)
Random ForestN/A55,883.222.833043,796
1D-CNN26,040123.7934.457229,022
Baseline DNN10,37273.8237.048326,992
DNN233956.9129.781333,578
Table 6. Comparison of model architecture.
Table 6. Comparison of model architecture.
ModelPreprocessingActivationArchitectureOptimization
Model AZ-ScoreReLuPlain DNNFixed LR
Model BRobustReLuPlain DNNFixed LR
Model CRobustSELUPlain DNNFixed LR
Model DRobustSELU+ResidualFixed LR
Model ERobustSELU+ResidualCosine
Table 7. Indicators of each model in the ablation experiment.
Table 7. Indicators of each model in the ablation experiment.
ModelAccuracyF1-ScoreInference Time (μs)
Model A0.9881230.98811018.43
Model B0.9922540.99224119.33
Model C0.99333100.99330119.69
Model D0.9951220.99511519.76
Model E0.9953410.9953319.76
Table 8. F1-score degradation of different models under varying levels of adversarial evasion noise.
Table 8. F1-score degradation of different models under varying levels of adversarial evasion noise.
Evasion Noise RatioRandom Forest (RF)Baseline CNNBaseline DNNDNN
0% (Clean Data)0.99030.89540.62290.9030
5% Evasion Noise0.93870.81230.57120.8876
10% Evasion Noise0.84120.72050.50880.8641
20% Evasion Noise0.62890.55140.41350.8293
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Cui, J.; Zhao, Y. Design of Network Traffic Analysis Models Based on Deep Neural Networks. Future Internet 2026, 18, 152. https://doi.org/10.3390/fi18030152

AMA Style

Cui J, Zhao Y. Design of Network Traffic Analysis Models Based on Deep Neural Networks. Future Internet. 2026; 18(3):152. https://doi.org/10.3390/fi18030152

Chicago/Turabian Style

Cui, Jiantao, and Yixiang Zhao. 2026. "Design of Network Traffic Analysis Models Based on Deep Neural Networks" Future Internet 18, no. 3: 152. https://doi.org/10.3390/fi18030152

APA Style

Cui, J., & Zhao, Y. (2026). Design of Network Traffic Analysis Models Based on Deep Neural Networks. Future Internet, 18(3), 152. https://doi.org/10.3390/fi18030152

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop