BHM-IDS: Behavior-Driven Hierarchy and Multi-Dataset Training for Cross-Dataset Generalization
Abstract
1. Introduction
- We propose a three-stage hierarchical IDS that incorporates behavior-based classification at the second stage.
- We introduce a multi-dataset training strategy that enriches CIC-IDS2017 with CIC-DDoS2019, enabling the model to learn from a broader and more diverse set of attacks.
- We perform a rigorous evaluation under both intra-dataset and cross-dataset settings using CSE-CIC-IDS2018 as an external test dataset, ensuring a realistic assessment of generalization capability.
- We conduct a detailed ablation study that isolates the respective roles of behavior-driven hierarchy and multi-dataset training, showing that their combination yields strong generalization gains.
2. Related Work
2.1. Architectural Advances in ML-Based IDS
2.2. Multi-Dataset Training for ML-Based IDSs
2.3. Generalization in ML-Based IDSs
2.4. Research Gaps and Motivation
3. BHM-IDS Framework: Design and Methodology
3.1. Hierarchical Design and Behavioral Grouping
3.2. Multi-Dataset Training and Generalization Assessment
3.3. Dataset Preparation
3.4. Stage-Specific Learning Pipeline
3.4.1. Stage 1
3.4.2. Stage 2
3.4.3. Stage 3
3.5. Learning Models and Hyperparameter Settings
3.6. Evaluation Strategies and Performance Metrics
4. Results
4.1. Intra-Dataset Validation
4.2. Cross-Dataset Validation
- Stage 1:The results reported in Table 13 present the cross-dataset comparison of XGBoost, CatBoost, and LightGBM. XGBoost provides the best overall performance, achieving the highest accuracy (0.9367). The main difference between the models appears in their ability to detect attack samples, where XGBoost reaches the highest Attack recall (0.7893), compared with LightGBM (0.6369) and CatBoost (0.5002). In contrast, all models achieve high recall for the Benign class, indicating that normal traffic is consistently well identified across the three classifiers. This behavior is consistent with the confusion matrix of XGBoost model in Figure 5a.To provide a more detailed analysis of Stage 1 performance, Table 14 reports the per-attack recall obtained by the XGBoost classifier. The results show that several attack categories are identified with near-perfect scores. In particular, recall reaches 1 for DoS SlowHttpTest, SSH-Patator, and FTP-Patator, while DoS Slowloris and DoS Hulk achieve recall values of 0.9845 and 0.9253, respectively. DDoS LOIC-HTTP is also detected with high reliability, reaching a recall of 0.9866. The results remain relatively lower for some other DDoS variants, Bot, and Web attack variants.
- Stage 2: The results of the three models are summarized in Table 15. LightGBM achieves the best overall performance, with the highest accuracy (0.9604) and balanced recall for both attack groups. In particular, it maintains high recall for Attack Group 1 (0.9612) and Attack Group 2 (0.9576). This result is further supported by the confusion matrix in Figure 5b. XGBoost provides competitive results but remains slightly below LightGBM, whereas CatBoost shows weaker generalization due to its limited ability to detect Attack Group 2.
- Stage 3: Under cross-dataset validation, Stage 3 classifiers exhibit excellent generalization performance, achieving near-perfect accuracy and consistently high recall. For the first classifier, dedicated to DoS/DDoS classification, CatBoost achieves the best overall performance, as shown in Table 16. It reaches an accuracy of 0.9999, with high recall values. This result is further supported by the confusion matrix in Figure 6a, which shows very little misclassification between the two categories.For the second classifier, which distinguishes between Web Attack, Bot, and Patator, XGBoost provides the most balanced results, as reported in Table 17. It attains an overall accuracy of 0.9994 while maintaining high recall across all classes. The corresponding confusion matrix in Figure 6b confirms this behavior.Finally, Table 18 reports the combined recall achieved by attack classification stages (Stage 2 and Stage 3). This metric serves as a crucial indicator of the system’s ability to correctly classify threats once they have passed the initial detection stage. It is calculated as the product of the class-specific recall values obtained from the best-performing models in Stage 2 and Stage 3. The results demonstrate high performance across all categories, with scores exceeding 0.96 for DoS and DDoS and remaining above 0.93 for Web attack, Bot, and Patator attacks. The final macro combined recall of 0.9535 further indicates that the proposed behavior-based design preserves strong attack classification capability under cross-dataset validation.
4.3. End-to-End Cross-Dataset Performance Evaluation of BHM-IDS
4.4. Ablation Study
4.4.1. Ablation Study Configurations
- Configuration 1 represents a conventional two-stage architecture composed of a binary classifier followed by a flat multi-class classifier, both trained exclusively on CIC-IDS2017. This configuration serves as the baseline architecture because it does not incorporate multi-dataset training or the behavior-based hierarchy introduced in BHM-IDS.
- Configuration 2 keeps the same structure used in Configuration 1, but changes the training setting by incorporating multi-dataset training. Since this configuration does not introduce behavior-based hierarchy, it allows the effect of multi-dataset training to be evaluated independently.
- Configuration 3 has the same structure as BHM-IDS, while restricting the training process to CIC-IDS2017 only. By relying on single-dataset training, this configuration isolates the impact of behavior-based hierarchy.
4.4.2. Ablation Study Results
- Configuration 1: The performance of the baseline architecture at Stage 1 is reported in Table 20. LightGBM achieves an overall accuracy of 0.8250, but its low attack recall of 0.4124 indicates limited effectiveness in detecting attack traffic despite strong benign classification performance.At the individual attack level, Stage 1 recall shows clear variability, as reported in Table 21. While some attacks such as DoS Slowhttptest and FTP-Patator are perfectly detected, others, including DDoS HOIC, DDoS LOIC-UDP, Web Attack, SSH-Patator, and Bot, show critically weak or zero detection.At Stage 2, the baseline architecture shows degraded attack classification performance, with an overall accuracy of 0.569, as reported in Table 22. While the Bot class achieves exceptional detection (Recall: 0.9991, Precision: 1.0), other classes exhibit remarkable drops in performance.
- Configuration 2:Under the multi-dataset training setting, Stage 2 of Configuration 2 achieves an overall accuracy of 0.7950 and a macro recall of 0.7674, as reported in Table 23. The configuration demonstrates relatively better performance for Patator, DoS, and DDoS, while Web Attack remains the main limitation, with a very low F1-score of 0.0158.
- Configuration 3:The results of the Stage 2 and Stage 3 classifiers are presented in Table 24, Table 25 and Table 26, respectively. At Stage 2, the classifier achieves a high overall accuracy of 0.9097; it is better at detecting Attack Group 2 and more precise when predicting Attack Group 1.At Stage 3, the DoS and DDoS classifier achieves a lower accuracy of 0.7013, mainly due to the reduced DDoS recall of 0.5804, despite its high precision of 0.9994. In contrast, the Web Attack, Bot, and Patator classifier achieves excellent performance, with an accuracy of 0.9994 and all class metrics exceeding 0.97.
5. Comparative Discussion
5.1. Impact of Multi-Dataset Training
5.2. Impact of Behavior-Driven Hierarchy
5.3. Synergistic Effect of Behavior-Driven Hierarchy and Multi-Dataset Training
5.4. Error Propagation Analysis
5.5. Comparison with State-of-the-Art Approaches
5.6. Limitations and Future Work
- Expanded Cross-Dataset Validation: Additional datasets will be incorporated into the evaluation pipeline to ensure broader and more robust cross-dataset validation.
- Class-Specific Failure Analysis: The factors underlying the weak detection performance observed for challenging attack categories, particularly Web Attack, will be investigated. To determine whether these limitations are associated with cross-dataset distribution shifts, the distributions of the most relevant Stage 1 features will be systematically analyzed and compared across the three adopted datasets. This analysis will combine quantitative measures, including the Kolmogorov–Smirnov statistic, Wasserstein distance, Jensen–Shannon divergence, Population Stability Index, and Maximum Mean Discrepancy, with visual techniques such as kernel-density estimation, empirical cumulative distribution plots, box and violin plots, and PCA/UMAP projections. These complementary analyses will help identify the causes of the observed class-specific performance degradation, provide deeper insight into dataset shift, and support the selection of more complementary and synergistic dataset combinations.
- Preprocessing Optimization: A broader range of preprocessing configurations will be explored, with a specific focus on stage-specific feature-selection methods, class-balancing strategies, and scaling techniques. Systematically optimizing these steps will help reduce dataset-specific bias and further enhance the robustness of BHM-IDS.
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Appendix A
| Packet Length Std | Subflow Fwd Packets | Bwd Packets/s |
| Packet Length Variance | Fwd IAT Max | Flow IAT Max |
| Avg Bwd Segment Size | PSH Flag Count | Fwd IAT Total |
| Max Packet Length | Flow IAT Std | Flow Packets/s |
| Bwd Packet Length Max | Bwd Header Length | Idle Min |
| Total Length of Bwd Packets | Idle Mean | Flow Duration |
| Bwd Packet Length Std | Fwd IAT Std | |
| Average Packet Size | Fwd Header Length | |
| Subflow Fwd Bytes | Init_Win_bytes_forward | |
| Fwd Packet Length Max | act_data_pkt_fwd | |
| Total Length of Fwd Packets | Flow Bytes/s | |
| Subflow Bwd Bytes | Fwd Packet Length Std | |
| Fwd Packet Length Mean | Fwd IAT Mean | |
| Avg Fwd Segment Size | Init_Win_bytes_backward | |
| Packet Length Mean | Flow IAT Mean | |
| Bwd Packet Length Mean | min_seg_size_forward | |
| Total Fwd Packets | Subflow Bwd Packets |
| Init_Win_bytes_backward | Fwd Header Length | Destination Port |
| Subflow Fwd Bytes | Subflow Fwd Packets | Idle Min |
| Bwd Packet Length Min | act_data_pkt_fwd | Avg Bwd Segment Size |
| Total Length of Fwd Packets | Flow IAT Max | Bwd Packet Length Mean |
| Fwd Packet Length Mean | Flow IAT Mean | Max Packet Length |
| Fwd Packet Length Max | Average Packet Size | Fwd IAT Min |
| Fwd Packet Length Std | Packet Length Mean | Flow Bytes/s |
| Bwd Header Length | Flow IAT Std | Packet Length Variance |
| Fwd IAT Mean | Bwd Packet Length Max | Packet Length Std |
| Avg Fwd Segment Size | Flow Packets/s | Idle Max |
| Fwd IAT Max | Init_Win_bytes_forward | Fwd IAT Std |
| Fwd IAT Total | Flow Duration | Total Backward Packets |
| min_seg_size_forward | Subflow Bwd Bytes | Flow IAT Min |
| Total Fwd Packets | Fwd Packets/s | Bwd Packet Length Std |
| Bwd Packets/s | Total Length of Bwd Packets | Bwd IAT Total |
| min_seg_size_forward | Fwd IAT Total | Packet Length Mean |
| Flow Packets/s | Avg Fwd Segment Size | Average Packet Size |
| Flow Duration | Fwd Packet Length Std | Idle Min |
| Flow IAT Mean | Fwd IAT Min | Subflow Bwd Bytes |
| Init_Win_bytes_backward | Flow IAT Min | Subflow Bwd Packets |
| Flow IAT Max | Fwd Header Length | Bwd Packet Length Max |
| Fwd Packets/s | Total Length of Fwd Packets | Max Packet Length |
| Init_Win_bytes_forward | Fwd Packet Length Mean | Packet Length Std |
| Fwd IAT Max | ACK Flag Count | Bwd IAT Max |
| Min Packet Length | Total Fwd Packets | Avg Bwd Segment Size |
| act_data_pkt_fwd | Bwd Header Length | |
| Flow Bytes/s | Idle Max | |
| Destination Port | Fwd Packet Length Max | |
| Subflow Fwd Bytes | Bwd Packet Length Min | |
| Fwd IAT Mean | Flow IAT Std |
| Destination Port | Average Packet Size | Fwd Packets/s |
| min_seg_size_forward | Packet Length Mean | Flow Duration |
| Max Packet Length | Subflow Bwd Bytes | Bwd Header Length |
| Packet Length Std | Flow IAT Max | Total Fwd Packets |
| Bwd IAT Min | Bwd Packet Length Max | Fwd Header Length |
| Init_Win_bytes_backward | Packet Length Variance | Avg Bwd Segment Size |
| Flow IAT Mean | Fwd Packet Length Mean | Fwd IAT Min |
| Bwd Packet Length Min | Fwd IAT Std | Bwd IAT Total |
| Fwd Packet Length Max | Bwd Packet Length Mean | Down/Up Ratio |
| Flow IAT Std | Flow Bytes/s | Subflow Fwd Packets |
| Subflow Fwd Bytes | Total Length of Bwd Packets | |
| Fwd IAT Max | Init_Win_bytes_forward | |
| Fwd Packet Length Std | Flow Packets/s | |
| Total Length of Fwd Packets | Avg Fwd Segment Size | |
| Fwd IAT Mean | Bwd Packets/s |
References
- AlNuaimi, B.K.; Singh, S.K.; Ren, S.; Budhwar, P.; Vorobyev, D. Mastering digital transformation: The nexus between leadership, agility, and digital strategy. J. Bus. Res. 2022, 145, 636–648. [Google Scholar] [CrossRef] [Scilit]
- Ly, B.; Ly, R.; Ma, S. Digital transformation and flexibility in public services: Knowledge, culture and digital infrastructures. J. Innov. Knowl. 2026, 13, 100947. [Google Scholar] [CrossRef] [Scilit]
- ENISA (European Union Agency for Cybersecurity). ENISA Threat Landscape 2022. Report, European Union Agency for Cybersecurity, 2022. Available online: https://www.enisa.europa.eu/publications/enisa-threat-landscape-2022 (accessed on 25 June 2025).
- Aljundi, I.; Rawashdeh, M.; Al-Fayoumi, M.; Al-Badarneh, A.; Al-Haija, Q.A. Protecting Critical National Infrastructures: An Overview of Cyberattacks and Countermeasures. In Proceedings of the Intelligent Sustainable Systems; Nagar, A.K., Jat, D.S., Mishra, D., Joshi, A., Eds.; Springer: Singapore, 2024; pp. 295–317. [Google Scholar] [CrossRef] [Scilit]
- Diana, L.; Dini, P.; Paolini, D. Overview on Intrusion Detection Systems for Computers Networking Security. Computers 2025, 14, 87. [Google Scholar] [CrossRef] [Scilit]
- Arnob, A.K.B.; Chowdhury, R.R.; Chaiti, N.A.; Saha, S.; Roy, A. A comprehensive systematic review of intrusion detection systems: Emerging techniques, challenges, and future research directions. J. Edge Comput. 2025, 4, 73–104. [Google Scholar] [CrossRef] [Scilit]
- Pinto, A.; Herrera, L.C.; Donoso, Y.; Gutierrez, J.A. Survey on Intrusion Detection Systems Based on Machine Learning Techniques for the Protection of Critical Infrastructure. Sensors 2023, 23, 2415. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hozouri, A.; Mirzaei, A.; Effatparvar, M. A comprehensive survey on intrusion detection systems with advances in machine learning, deep learning and emerging cybersecurity challenges. Discov. Artif. Intell. 2025, 5, 314. [Google Scholar] [CrossRef] [Scilit]
- Ahmad, Z.; Shahid Khan, A.; Wai Shiang, C.; Abdullah, J.; Ahmad, F. Network intrusion detection system: A systematic study of machine learning and deep learning approaches. Trans. Emerg. Telecommun. Technol. 2021, 32, e4150. [Google Scholar] [CrossRef] [Scilit]
- Yu, H.; Zhang, W.; Kang, C.; Xue, Y. A feature selection algorithm for intrusion detection system based on the enhanced heuristic optimizer. Expert Syst. Appl. 2025, 265, 125860. [Google Scholar] [CrossRef] [Scilit]
- Qi, Z.; Fei, J.; Wang, J.; Li, X. An Intrusion Detection Feature Selection Method Based on Improved Mutual Information. In Proceedings of the 2023 IEEE 6th Information Technology, Networking, Electronic and Automation Control Conference (ITNEC); IEEE: Piscataway, NJ, USA, 2023; Volume 6, pp. 1584–1590. [Google Scholar] [CrossRef] [Scilit]
- Rahma, F.; Rachmadi, R.F.; Pratomo, B.A.; Purnomo, M.H. Assessing the Effectiveness of Oversampling and Undersampling Techniques for Intrusion Detection on an Imbalanced Dataset. In Proceedings of the 2023 IEEE Industrial Electronics and Applications Conference (IEACon); IEEE: Piscataway, NJ, USA, 2023; pp. 92–97. [Google Scholar] [CrossRef] [Scilit]
- Othman, T.S.; Abdullah, S.M. Machine Learning Techniques Evaluation with SMOTE on IoT-23 Dataset. In Proceedings of the 2023 9th International Engineering Conference on Sustainable Technology and Development (IEC); IEEE: Piscataway, NJ, USA, 2023; pp. 7–13. [Google Scholar] [CrossRef] [Scilit]
- Khan, F.A.; Shah, A.A.; Alshammry, N.; Saif, S.; Khan, W.; Malik, M.O.; Ullah, Z. Balanced Multi-Class Network Intrusion Detection Using Machine Learning. IEEE Access 2024, 12, 178222–178236. [Google Scholar] [CrossRef] [Scilit]
- Lu, C.; Cao, Y.; Wang, Z. Research on Intrusion Detection Based on an Enhanced Random Forest Algorithm. Appl. Sci. 2024, 14, 714. [Google Scholar] [CrossRef] [Scilit]
- Srivastav, N.; Singh, R. An Optimized Machine Learning Based Network Intrusion Detection Systems for Identification of Low-Occurrence Attacks. SN Comput. Sci. 2025, 6, 820. [Google Scholar] [CrossRef] [Scilit]
- Cao, B.; Li, C.; Song, Y.; Qin, Y.; Chen, C. Network Intrusion Detection Model Based on CNN and GRU. Appl. Sci. 2022, 12, 4184. [Google Scholar] [CrossRef] [Scilit]
- Qazi, E.U.H.; Almorjan, A.; Zia, T. A One-Dimensional Convolutional Neural Network (1D-CNN) Based Deep Learning System for Network Intrusion Detection. Appl. Sci. 2022, 12, 7986. [Google Scholar] [CrossRef] [Scilit]
- Verkerken, M.; D’hooge, L.; Sudyana, D.; Lin, Y.D.; Wauters, T.; Volckaert, B.; De Turck, F. A Novel Multi-Stage Approach for Hierarchical Intrusion Detection. IEEE Trans. Netw. Serv. Manag. 2023, 20, 3915–3929. [Google Scholar] [CrossRef] [Scilit]
- Alin, F.; Chemchem, A.; Nolot, F.; Flauzac, O.; Krajecki, M. Towards a Hierarchical Deep Learning Approach for Intrusion Detection. In Proceedings of the Machine Learning for Networking; Boumerdassi, S., Renault, É., Mühlethaler, P., Eds.; Springer: Cham, Switzerland, 2020; pp. 15–27. [Google Scholar] [CrossRef] [Scilit]
- Hewapathirana, I.U. A Comparative Study of Two-Stage Intrusion Detection Using Modern Machine Learning Approaches on the CSE-CIC-IDS2018 Dataset. Knowledge 2025, 5, 6. [Google Scholar] [CrossRef] [Scilit]
- Uddin, M.A.; Aryal, S.; Bouadjenek, M.R.; Al-Hawawreh, M.; Talukder, M.A. Hierarchical classification for intrusion detection system: Effective design and empirical analysis. Ad Hoc Netw. 2025, 178, 103982. [Google Scholar] [CrossRef] [Scilit]
- Mihai, I.C.; Pruna, S.; Barbu, I.D. Cyber Kill Chain Analysis. Int. J. Inf. Secur. Cybercrime 2014, 3, 37–42. [Google Scholar] [CrossRef] [Scilit]
- Kim, Y.; Kim, J.; Kim, D. Hi-MLIC: Hierarchical Multilayer Lightweight Intrusion Classification for Various Intrusion Scenarios. IEEE Access 2024, 12, 120098–120115. [Google Scholar] [CrossRef] [Scilit]
- Blank, R.; Gallagher, P. Guide for Conducting Risk Assessments; Technical Report NIST Special Publication 800-30 Revision 1; National Institute of Standards and Technology (NIST): Gaithersburg, MD, USA, 2012. [Google Scholar] [CrossRef] [Scilit]
- Sarıkaya, A.; Kılıç, B.G. A Class-Specific Intrusion Detection Model: Hierarchical Multi-class IDS Model. SN Comput. Sci. 2020, 1, 202. [Google Scholar] [CrossRef] [Scilit]
- Uddin, M.A.; Aryal, S.; Bouadjenek, M.R.; Al-Hawawreh, M.; Talukder, M.A. A dual-tier adaptive one-class classification IDS for emerging cyberthreats. Comput. Commun. 2025, 229, 108006. [Google Scholar] [CrossRef] [Scilit]
- Mohd, N.; Singh, A.; Bhadauria, H.S. Intrusion Detection System Based on Hybrid Hierarchical Classifiers. Wirel. Pers. Commun. 2021, 121, 659–686. [Google Scholar] [CrossRef] [Scilit]
- ElDahshan, K.A.; AlHabshy, A.A.; Hameed, B.I. Meta-Heuristic Optimization Algorithm-Based Hierarchical Intrusion Detection System. Computers 2022, 11, 170. [Google Scholar] [CrossRef] [Scilit]
- Kilichev, D.; Kim, W. Hyperparameter Optimization for 1D-CNN-Based Network Intrusion Detection Using GA and PSO. Mathematics 2023, 11, 3724. [Google Scholar] [CrossRef] [Scilit]
- Sahu, P.; Vyas, O.P.; Barnwal, R.; Singla, A.; Priyanshu. Enhancing Industrial IoT Intrusion Detection with Hyperparameter Optimization. In Proceedings of the 2024 15th International Conference on Computing Communication and Networking Technologies (ICCCNT); IEEE: Piscataway, NJ, USA, 2024; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Hairab, B.I.; Said Elsayed, M.; Jurcut, A.D.; Azer, M.A. Anomaly Detection Based on CNN and Regularization Techniques Against Zero-Day Attacks in IoT Networks. IEEE Access 2022, 10, 98427–98440. [Google Scholar] [CrossRef] [Scilit]
- Tavallaee, M.; Bagheri, E.; Lu, W.; Ghorbani, A.A. A detailed analysis of the KDD CUP 99 data set. In Proceedings of the IEEE Symposium on Computational Intelligence for Security and Defense Applications (CISDA); IEEE: Piscataway, NJ, USA, 2009; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Choudhary, S.; Kesswani, N. Analysis of KDD-Cup’99, NSL-KDD and UNSW-NB15 Datasets using Deep Learning in IoT. Procedia Comput. Sci. 2020, 167, 1561–1573. [Google Scholar] [CrossRef] [Scilit]
- Sharafaldin, I.; Lashkari, A.H.; Ghorbani, A.A. Toward Generating a New Intrusion Detection Dataset and Intrusion Traffic Characterization. In Proceedings of the 4th International Conference on Information Systems Security and Privacy (ICISSP); Springer: Cham, Switzerland, 2018; pp. 108–116. [Google Scholar] [CrossRef] [Scilit]
- Magán-Carrión, R.; Urda, D.; Diaz-Cano, I.; Dorronsoro, B. Improving the Reliability of Network Intrusion Detection Systems Through Dataset Integration. IEEE Trans. Emerg. Top. Comput. 2022, 10, 1717–1732. [Google Scholar] [CrossRef] [Scilit]
- Iwanowski, M.; Olszewski, D.; Graniszewski, W.; Krupski, J.; Pelc, F. The Choice of Training Data and the Generalizability of Machine Learning Models for Network Intrusion Detection Systems. Appl. Sci. 2025, 15, 8466. [Google Scholar] [CrossRef] [Scilit]
- Cantone, M.; Marrocco, C.; Bria, A. Machine Learning in Network Intrusion Detection: A Cross-Dataset Generalization Study. IEEE Access 2024, 12, 144489–144508. [Google Scholar] [CrossRef] [Scilit]
- Verkerken, M.; D’hooge, L.; Wauters, T.; Volckaert, B.; De Turck, F. Towards Model Generalization for Intrusion Detection: Unsupervised Machine Learning Techniques. J. Netw. Syst. Manag. 2021, 30, 12. [Google Scholar] [CrossRef] [Scilit]
- Xu, C.; Li, D.; Liu, Z.; Yang, J.; Shen, Q.; Tong, N. Few-shot network intrusion detection method based on multi-domain fusion and cross-attention. PLoS ONE 2025, 20, e0327161. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Layeghy, S.; Portmann, M. Explainable Cross-domain Evaluation of ML-based Network Intrusion Detection Systems. Comput. Electr. Eng. 2023, 108, 108692. [Google Scholar] [CrossRef] [Scilit]
- Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD); ACM: Singapore, 2016; pp. 785–794. [Google Scholar] [CrossRef] [Scilit]
- Prokhorenkova, L.; Gusev, G.; Vorobev, A.; Dorogush, A.V.; Gulin, A. CatBoost: Unbiased Boosting with Categorical Features. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2018; Volume 31, pp. 6639–6649. Available online: https://proceedings.neurips.cc/paper/2018/hash/14491b756b3a51daac41c24863285549-Abstract.html (accessed on 20 December 2025).
- Ke, G.; Meng, Q.; Finley, T.; Wang, T.; Chen, W.; Ma, W.; Ye, Q.; Liu, T.Y. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30, Available online: https://proceedings.neurips.cc/paper/2017/hash/6449f44a102fde848669bdd9eb6b76fa-Abstract.html (accessed on 26 December 2025).
- Bergstra, J.; Bengio, Y. Random search for hyper-parameter optimization. J. Mach. Learn. Res. 2012, 13, 281–305. [Google Scholar]
- Akiba, T.; Sano, S.; Yanase, T.; Ohta, T.; Koyama, M. Optuna: A Next-generation Hyperparameter Optimization Framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD); ACM: Singapore, 2019; pp. 2623–2631. [Google Scholar] [CrossRef] [Scilit]










| Study | Stages | Attack Grouping Technique | No. of Groups | Datasets | ML/DL Models |
|---|---|---|---|---|---|
| [19] | 2 | – | – | CIC-IDS2017 | Autoencoder, OCSVM, Random Forest |
| [20] | 2 | – | – | KDD-99, NSL-KDD, UNSW-NB15, CIC-IDS2017 | Naïve Bayes, Decision Tree, MLP, Logistic Regression, KNN |
| [21] | 2 | – | – | CSE-CIC-IDS2018 | Autoencoder, MLP, Logistic Regression, Random Forest |
| [22] | 3 | Cyber kill chain-based | Dataset dependent | CIC-DoS2017, CIC-DDoS2019, NSL-KDD, ToN-IoT | Random Forest, Decision Tree, Logistic Regression, MLP, Gaussian NB, Extra Trees |
| [24] | 3 | NIST-Based | 4 | CIC-IDS2017, UNSW-NB15 | Random Forest, MLP, Decision Tree, KNN |
| [26] | 3 | Error-driven approach | 2 | UNSW-NB15 | Random Forest |
| Study | Combined Datasets | Architectural Design | ML/DL Models | Cross-Dataset Validation |
|---|---|---|---|---|
| [36] | UGR-16, UNSW-NB15, NSL-KDD.BH | Flat | Logistic Regression, Random Forest. | Evaluation mainly on the new and the constituent datasets; no systematic testing on an additional independent dataset |
| [24] | CIC-IDS2017, UNSW-NB15 | Hierarchical | Random Forest, Multilayer Perceptron, Decision Tree, K-Nearest Neighbor. | Evaluation confined to the new dataset; no separate unseen dataset used |
| Study | Datasets | Multi Class Classification | Binary Classification | ML/DL Models | Best Train–Test Pair |
|---|---|---|---|---|---|
| [37] | UNSW-NB15, CIC-CSE-IDS2018, BoT-IoT, ToN-IoT | ✓ | X | Decision Tree, Random Forest | CSE-IDS2018–UNSW-NB15 |
| [38] | CIC-IDS2017, CSE-CIC-IDS2018, Lycos-IDS2017, Lycos-Unicas-IDS2018 | X | ✓ | Linear Discriminant Analysis, Decision Tree, Random Forest, XGBoost | CIC-IDS2018–
CIC-IDS2017 |
| [39] | CIC-IDS2017, CSE-CIC-IDS2018 | X | ✓ | Isolation Forest, Autoencoder, OCSVM | CIC-IDS2018– CIC-IDS2017 |
| [40] | CIC-IDS2017, CSE-CIC-IDS2018 | ✓ | X | Multi-Layer Perceptron | CIC-IDS2018– CIC-IDS2017 |
| [41] | UNSW-NB15, CIC-CSE-IDS2018, ToN-IoT, BoT-IoT | X | ✓ | Decision Tree, Random Forest, Extra Trees | UNSW-NB15– BoT-IoT |
| Criterion | Cyber Kill Chain-Based | NIST-Based | Behavioral-Based |
|---|---|---|---|
| Grouping Principle | Attacks are grouped according to their dominant role in the attack lifecycle. | Attacks are grouped according to operational impact. | Attacks are grouped according to observable traffic behavior. |
| Nature of Grouping | Semantic and stage based | Semantic and impact based | Data driven and behavior based |
| Relation to Feature Space | Indirect. | Indirect. | Direct: grouping is aligned with measurable traffic features. |
| Dependence on Contextual Interpretation | High: requires interpretation of the attack’s role within a broader scenario | High: depends on how attack categories are defined and labeled | Low: grouping is driven by traffic patterns, without requiring interpretation of attack intent or context |
| Consistency Across Datasets | Limited: the same attack may be assigned differently depending on scenario interpretation | Limited: category definitions and label granularity may vary across datasets | High: behavioral patterns are more stable across datasets than semantic labels |
| Suitability for ML | Moderate | Moderate | High |
| Expected Cross-Dataset Transferability | Moderate | Moderate | High |
| Traffic Type | CIC-IDS2017 | % | CSE-CIC-IDS2018 | % | New Label |
|---|---|---|---|---|---|
| Benign | 2,271,320 | 85.099% | 13,390,249 | 83.812% | Benign |
| DoS GoldenEye | 10,293 | 0.386% | 41,508 | 0.260% | DoS |
| DoS Slowloris | 5796 | 0.217% | 10,990 | 0.069% | |
| DoS Hulk | 230,124 | 8.622% | 461,912 | 2.891% | |
| DoS Slowhttptest | 5499 | 0.206% | 139,890 | 0.876% | |
| DDoS LOIC-HTTP | 128,025 | 4.797% | 576,191 | 3.606% | DDoS |
| DDoS LOIC-UDP | 0 | 0.000% | 1730 | 0.011% | |
| DDoS HOIC | 0 | 0.000% | 686,012 | 4.294% | |
| SSH-Patator | 5897 | 0.221% | 187,589 | 1.174% | Patator |
| FTP-Patator | 7935 | 0.297% | 193,354 | 1.210% | |
| Botnet | 1956 | 0.073% | 286,191 | 1.791% | Bot |
| Web Attack—Brute Force | 1507 | 0.056% | 611 | 0.004% | Web Attack |
| Web Attack—XSS | 652 | 0.024% | 230 | 0.001% | |
| Web Attack—SQL Injection | 21 | 0.001% | 87 | 0.001% | |
| Total | 2,669,025 | 100% | 15,976,544 | 100% | – |
| Traffic Type | Instances | % | New Label |
|---|---|---|---|
| Benign | 2,273,097 | 1.8978% | Benign |
| DDoS-DNS | 5,071,011 | 4.2338% | DDoS |
| DDoS-LDAP | 21,799,830 | 18.2009% | |
| DDoS-MSSQL | 4,522,492 | 3.7759% | |
| DDoS-NTP | 1,202,642 | 1.0041% | |
| DDoS-SNMP | 5,159,870 | 4.3080% | |
| DDoS-NetBIOS | 4,093,279 | 3.4175% | |
| DDoS-SSDP | 26,106,511 | 21.7966% | |
| DDoS-SYN | 15,822,889 | 13.2107% | |
| DDoS-TFTP | 2,008,258 | 1.6767% | |
| DDoS-UDP | 31,346,455 | 26.1715% | |
| DDoS-UDPLag | 366,461 | 0.3060% | |
| DDoS-WebDDoS | 439 | 0.0004% | |
| Total | 119,773,234 | 100% | – |
| Stage | Top Ten Selected Features |
|---|---|
| Stage 1 | |
| Benign/Attack | |
| |
| Stage 2 | |
| Attack Group 1/Attack Group 2 | |
| |
| Stage 3 (Classifier 1) | |
| DoS/DDoS | |
| |
| Stage 3 (Classifier 2) | |
| Bot/Patator/Web Attack | |
|
| ML Model | Hyperparameter | Search Space | Stage 1 | Stage 2 | Stage 3 | |
|---|---|---|---|---|---|---|
| Classifier 1 | Classifier 2 | |||||
| XGBoost | n_estimators | 300 | 100 | 300 | 300 | |
| max_depth | 10 | 10 | 10 | 6 | ||
| learning_rate | 0.05 | 0.02 | 0.05 | 0.05 | ||
| subsample | 0.9 | 0.9 | 0.9 | 0.7 | ||
| min_child_weight | 5 | 5 | 5 | 5 | ||
| CatBoost | iterations | 500 | 250 | 500 | 500 | |
| depth | 6 | 10 | 10 | 10 | ||
| learning_rate | 0.01 | 0.01 | 0.01 | 0.01 | ||
| rsm | 0.7 | 0.7 | 0.7 | 0.9 | ||
| l2_leaf_reg | 3 | 3 | 5 | 3 | ||
| LightGBM | n_estimators | 500 | 500 | 500 | 300 | |
| max_depth | 6 | 3 | 6 | 3 | ||
| learning_rate | 0.05 | 0.01 | 0.01 | 0.01 | ||
| colsample_bytree | 0.7 | 0.7 | 0.9 | 0.7 | ||
| num_leaves | 63 | 43 | 63 | 63 | ||
| ML Model | Class | Acc | Pre | Rec | F1 |
|---|---|---|---|---|---|
| XGBoost | Benign | 0.9950 | 0.9947 | 0.9954 | 0.9950 |
| Attack | 0.9954 | 0.9947 | 0.9950 | ||
| CatBoost | Benign | 0.9977 | 0.9985 | 0.9970 | 0.9977 |
| Attack | 0.9970 | 0.9985 | 0.9977 | ||
| LightGBM | Benign | 0.9987 | 0.9992 | 0.9983 | 0.9987 |
| Attack | 0.9983 | 0.9992 | 0.9987 |
| ML Model | Attack Group | Acc | Pre | Rec | F1 |
|---|---|---|---|---|---|
| XGBoost | Attack Group 1 | 0.9993 | 0.9999 | 0.9987 | 0.9993 |
| Attack Group 2 | 0.9987 | 0.9999 | 0.9993 | ||
| CatBoost | Attack Group 1 | 0.9999 | 1 | 0.9999 | 0.9999 |
| Attack Group 2 | 0.9999 | 1 | 0.9999 | ||
| LightGBM | Attack Group 1 | 0.9999 | 1 | 0.9999 | 0.9999 |
| Attack Group 2 | 0.9999 | 1 | 0.9999 |
| ML Model | Attack Type | Acc | Pre | Rec | F1 |
|---|---|---|---|---|---|
| XGBoost | DoS | 0.9997 | 0.9994 | 1 | 0.9997 |
| DDoS | 1 | 0.9994 | 0.9997 | ||
| CatBoost | DoS | 0.9999 | 0.9999 | 0.9999 | 0.9999 |
| DDoS | 0.9999 | 0.9999 | 0.9999 | ||
| LightGBM | DoS | 0.9998 | 0.9997 | 1 | 0.9998 |
| DDoS | 1 | 0.9997 | 0.9998 |
| ML Model | Attack Type | Acc | Pre | Rec | F1 |
|---|---|---|---|---|---|
| XGBoost | Web attack | 1 | 1 | 1 | 1 |
| Bot | 1 | 1 | 1 | ||
| Patator | 1 | 1 | 1 | ||
| CatBoost | Web attack | 1 | 1 | 1 | 1 |
| Bot | 1 | 1 | 1 | ||
| Patator | 1 | 1 | 1 | ||
| LightGBM | Web attack | 1 | 1 | 1 | 1 |
| Bot | 1 | 1 | 1 | ||
| Patator | 1 | 1 | 1 |
| ML Model | Class | Acc | Pre | Rec | F1 |
|---|---|---|---|---|---|
| CatBoost | Benign | 0.8531 | 0.8287 | 0.9977 | 0.9061 |
| Attack | 0.9890 | 0.5002 | 0.6647 | ||
| LightGBM | Benign | 0.8944 | 0.8690 | 0.9977 | 0.9294 |
| Attack | 0.9914 | 0.6369 | 0.7743 | ||
| XGBoost | Benign | 0.9367 | 0.9195 | 0.9977 | 0.9570 |
| Attack | 0.9931 | 0.7893 | 0.8795 | ||
| Macro avg | 0.9563 | 0.8935 | 0.9183 |
| Attack Type | Recall Value |
|---|---|
| DoS SlowHttpTest | 1 |
| DoS Hulk | 0.9253 |
| DoS GoldenEye | 0.7011 |
| DoS SlowLoris | 0.9845 |
| SSH-Patator | 1 |
| FTP-Patator | 1 |
| DDoS HOIC | 0.6126 |
| DDoS LOIC-UDP | 0.3094 |
| DDoS LOIC-HTTP | 0.9866 |
| Web Att BruteForce | 0.3044 |
| Web Att XSS | 0.2522 |
| Web Att SQL Injection | 0.1379 |
| Bot | 0.4908 |
| ML Model | Class | Acc | Pre | Rec | F1 |
|---|---|---|---|---|---|
| XGBoost | Attack Group 1 | 0.9541 | 0.9858 | 0.9564 | 0.9709 |
| Attack Group 2 | 0.8447 | 0.9450 | 0.8920 | ||
| CatBoost | Attack Group 1 | 0.8621 | 0.8679 | 0.9562 | 0.9188 |
| Attack Group 2 | 0.8111 | 0.4080 | 0.5429 | ||
| LightGBM | Attack Group 1 | 0.9604 | 0.9890 | 0.9612 | 0.9749 |
| Attack Group 2 | 0.8609 | 0.9576 | 0.9067 | ||
| Macro avg | 0.9250 | 0.9594 | 0.9408 |
| Model | Metric | DoS | DDoS | Accuracy | Macro |
|---|---|---|---|---|---|
| Average | |||||
| XGBoost | Precision | 1.0000 | 0.9652 | 0.9743 | 0.9826 |
| Recall | 0.9111 | 1.0000 | 0.9556 | ||
| F1-Score | 0.9535 | 0.9823 | 0.9679 | ||
| LightGBM | Precision | 1.0000 | 0.9846 | 0.9889 | 0.9923 |
| Recall | 0.9615 | 1.0000 | 0.9808 | ||
| F1-Score | 0.9804 | 0.9924 | 0.9864 | ||
| CatBoost | Precision | 1.0000 | 0.9999 | 0.9999 | 1.0000 |
| Recall | 0.9998 | 1.0000 | 0.9999 | ||
| F1-Score | 0.9999 | 1.0000 | 1.0000 |
| Model | Metric | Web | Bot | Patator | Average | Macro |
|---|---|---|---|---|---|---|
| Attack | Accuracy | |||||
| XGBoost | Precision | 1.0000 | 1.0000 | 0.9982 | 0.9994 | 0.9994 |
| Recall | 0.9720 | 0.9991 | 1.0000 | 0.9904 | ||
| F1-Score | 0.9858 | 0.9995 | 0.9991 | 0.9948 | ||
| LightGBM | Precision | 0.9961 | 0.9995 | 0.9982 | 0.9991 | 0.9979 |
| Recall | 0.8319 | 0.9990 | 1.0000 | 0.9436 | ||
| F1-Score | 0.9066 | 0.9993 | 0.9991 | 0.9683 | ||
| CatBoost | Precision | 1.0000 | 1.0000 | 0.9982 | 0.9992 | 0.9994 |
| Recall | 0.9666 | 0.9990 | 0.9991 | 0.9882 | ||
| F1-Score | 0.9830 | 0.9995 | 0.9991 | 0.9939 |
| Attack Class | Web Attack | DoS | DDoS | Bot | Patator |
|---|---|---|---|---|---|
| Combined Recall | 0.9309 | 0.9610 | 0.9612 | 0.9567 | 0.9576 |
| Macro Combined Recall | 0.9535 | ||||
| Level | Class/Metric | Precision | Recall | F1-Score | Support |
|---|---|---|---|---|---|
| Class-level | Benign | 0.9195 | 0.9977 | 0.9570 | 5,292,116 |
| Patator | 0.8624 | 1.0000 | 0.9261 | 156,668 | |
| DoS | 0.9848 | 0.8480 | 0.9113 | 506,075 | |
| DDoS | 0.9935 | 0.7610 | 0.8618 | 1,246,382 | |
| Bot | 0.8342 | 0.4453 | 0.5807 | 282,310 | |
| Web Attack | 0.0129 | 0.2705 | 0.0247 | 928 | |
| Aggregate | Macro average | 0.7679 | 0.7204 | 0.7103 | 7,484,479 |
| Weighted average | 0.9317 | 0.9273 | 0.9231 | 7,484,479 | |
| Overall | Accuracy | 0.9273 | 7,484,479 |
| Model | Class | Accuracy | Precision | Recall | F1 Score |
|---|---|---|---|---|---|
| LightGBM | Benign | 0.8250 | 0.8036 | 0.9959 | 0.8895 |
| Attack | 0.9767 | 0.4124 | 0.5799 |
| Attack Type | Recall Value |
|---|---|
| DoS SlowHttpTest | 1 |
| DoS Hulk | 0.9519 |
| DoS GoldenEye | 0.6780 |
| DoS SlowLoris | 0.9845 |
| SSH-Patator | 0.2003 |
| FTP-Patator | 1 |
| DDoS HOIC | 0 |
| DDoS LOIC-UDP | 0 |
| DDoS LOIC-HTTP | 0.5022 |
| Web Att BruteForce | 0.0016 |
| Web Att XSS | 0 |
| Web Att SQL Injection | 0.0345 |
| Bot | 0.2847 |
| Class | Web Attack | DoS | DDoS | Bot | Patator | Accuracy |
|---|---|---|---|---|---|---|
| Precision | 0.0087 | 0.3712 | 0.9996 | 1 | 0.9976 | 0.5698 |
| Recall | 0.5894 | 0.8555 | 0.3341 | 0.9991 | 0.7488 | |
| F1-Score | 0.0171 | 0.5178 | 0.5008 | 0.9995 | 0.8555 | |
| Macro recall | 0.705 | |||||
| Class | Web Attack | DoS | DDoS | Bot | Patator | Accuracy |
|---|---|---|---|---|---|---|
| Precision | 0.0080 | 0.6081 | 0.9276 | 1.0000 | 0.8839 | 0.7950 |
| Recall | 0.4735 | 0.9201 | 0.7316 | 0.7545 | 0.9574 | |
| F1-Score | 0.0158 | 0.7323 | 0.8180 | 0.8601 | 0.9192 | |
| Macro recall | 0.7674 | |||||
| Model | Metric | Attack Group 1 | Attack Group 2 | Accuracy |
|---|---|---|---|---|
| XGBoost | Precision | 0.9980 | 0.6917 | 0.9097 |
| Recall | 0.8889 | 0.9927 | ||
| F1-Score | 0.9403 | 0.8153 |
| Model | Metric | DoS | DDoS | Accuracy |
|---|---|---|---|---|
| XGBoost | Precision | 0.4916 | 0.9994 | 0.7013 |
| Recall | 0.9991 | 0.5804 | ||
| F1-Score | 0.6589 | 0.7343 |
| Model | Metric | Web Attack | Bot | Patator | Accuracy |
|---|---|---|---|---|---|
| XGBoost | Precision | 1 | 1 | 0.9982 | 0.9994 |
| Recall | 0.9720 | 0.9991 | 1 | ||
| F1-Score | 0.9858 | 0.9995 | 0.9991 |
| Attack | Web Attack | DoS | DDoS | Bot | Patator |
|---|---|---|---|---|---|
| Combined recall(Stage2×Stage3) | 0.9649 | 0.8880 | 0.5159 | 0.9918 | 0.9927 |
| Macro combined recall | 0.870 | ||||
| Characteristics | [38] | [39] | [40] | BHM-IDS |
|---|---|---|---|---|
| Training Dataset | CIC-IDS2017 | CIC-IDS2017 | CIC-IDS2017 | CIC-IDS2017 + CIC-DDoS2019 |
| Testing Dataset | CSE-CIC-IDS2018 | CSE-CIC-IDS2018 | CSE-CIC-IDS2018 | CSE-CIC-IDS2018 |
| Binary Classification | ✓ | ✓ | × | ✓ |
| Multi-class Classification | × | × | ✓ | ✓ |
| Evaluation Metrics Shared with BHM-IDS | F1-score | Recall, F1-score, Precision | Recall | _ |
| Metric | [39] | BHM-IDS |
|---|---|---|
| Precision | 0.7748 | 0.9931 |
| Recall | 0.9305 | 0.7893 |
| F1-Score | 0.8455 | 0.8795 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zekiouk, M.; Bencheikh Lehocine, M.; Bouzeraa, Y.; Bouanane, A.; Hristov, G.; Zahariev, P. BHM-IDS: Behavior-Driven Hierarchy and Multi-Dataset Training for Cross-Dataset Generalization. Appl. Sci. 2026, 16, 7885. https://doi.org/10.3390/app16167885
Zekiouk M, Bencheikh Lehocine M, Bouzeraa Y, Bouanane A, Hristov G, Zahariev P. BHM-IDS: Behavior-Driven Hierarchy and Multi-Dataset Training for Cross-Dataset Generalization. Applied Sciences. 2026; 16(16):7885. https://doi.org/10.3390/app16167885
Chicago/Turabian StyleZekiouk, Mounira, Madjed Bencheikh Lehocine, Yehya Bouzeraa, Ahlam Bouanane, Georgi Hristov, and Plamen Zahariev. 2026. "BHM-IDS: Behavior-Driven Hierarchy and Multi-Dataset Training for Cross-Dataset Generalization" Applied Sciences 16, no. 16: 7885. https://doi.org/10.3390/app16167885
APA StyleZekiouk, M., Bencheikh Lehocine, M., Bouzeraa, Y., Bouanane, A., Hristov, G., & Zahariev, P. (2026). BHM-IDS: Behavior-Driven Hierarchy and Multi-Dataset Training for Cross-Dataset Generalization. Applied Sciences, 16(16), 7885. https://doi.org/10.3390/app16167885

