Skip to Content
Future InternetFuture Internet
  • Article
  • Open Access

30 September 2026

19 Pages

A Task-Specific Autoencoder–CatBoost Framework for Intrusion Detection in Imbalanced IIoT

,
,
and
1
Department of Computer Science, University of Batna 2, Batna 05078, Algeria
2
Department of Computer Science, University 8 Mai 1945 Guelma, BP 401, Guelma 24000, Algeria
*
Author to whom correspondence should be addressed.

Abstract

Industrial Internet of Things (IIoT) intrusion detection remains challenging because of class imbalance, rare attacks, and heterogeneous evaluation protocols. This study proposes a modular intrusion-detection framework based on task-specific autoencoder representation learning and downstream CatBoost classification. The binary branch learns a 16-dimensional representation from benign traffic and combines it with reconstruction error, whereas the multiclass branch uses a separate 24-dimensional representation learned from malicious traffic. On ML-EdgeIIoT, the extended five-fold group-aware binary evaluation achieved ROC-AUC = 0.9925, AP = 0.9981, and F1 = 0.9829, with a benign FPR of 0.1113. A reconstruction-error ablation increased F1 from 0.9800 to 0.9829 and reduced FPR from 0.1311 to 0.1113. The complete 14-class evaluation achieved macro-F1 = 0.6175 and macro-AUC = 0.9549, while the 15-label cascade reached strict end-to-end attack-type success = 0.6584. External validation on WUSTL-IIoT-2021 yielded a five-label macro-F1 of 0.9408 and strict success of 0.9993. These results support the task-specific modular design while highlighting the remaining difficulty of fine-grained attack attribution and false-alarm reduction.

1. Introduction

The Industrial Internet of Things (IIoT) integrates interconnected devices, sensors, actuators, cyber-physical systems, and edge-computing components to support industrial monitoring and automation [1]. This connectivity also increases exposure to cyber threats, making intrusion detection systems (IDSs) important for protecting industrial networks. Because these environments often require continuous availability, IDS-based monitoring is particularly relevant [2]. Machine learning (ML) and deep learning (DL) have therefore been widely investigated for IoT and IIoT intrusion detection [3,4,5].
Despite recent progress, IDS evaluation in IoT and IIoT remains sensitive to differences in datasets, protocols, and reporting practices, which limit reproducibility and direct comparison across studies [6,7]. Accuracy alone is also insufficient in imbalanced multiclass settings, where dominant classes can mask poor performance on rare attacks [8,9].
Diverse benchmark datasets are therefore important for evaluating IDSs under heterogeneous traffic conditions. CICIoT2023, for example, contains 33 attacks grouped into seven categories [7]. This study uses Edge-IIoTset as the primary benchmark because it was designed for IoT/IIoT cybersecurity research and contains benign and malicious traffic across heterogeneous attack scenarios [10].
Class imbalance remains a major challenge in IIoT intrusion detection because dominant traffic patterns can reduce sensitivity to rare attacks [11]. High-dimensional traffic data may also contain redundant or weakly informative features, increasing computational cost and overfitting risk. Feature extraction and dimensionality reduction are therefore important for learning compact IDS representations [9,12].
Recent IIoT IDS research includes lightweight tabular, deep-learning, ensemble, and hybrid approaches, including ensemble methods for IIoT-MEC environments [13]. Autoencoder-based representation learning has also been combined with supervised classifiers [14,15,16,17], while CatBoost has been investigated in related IoT and IIoT IDS settings [18,19].
Therefore, the contribution of this study does not lie in simply combining an autoencoder with a supervised classifier. Among the closest studies reviewed, separate task-specific autoencoders trained on benign and malicious traffic for binary detection and attack attribution, respectively, have not been explicitly evaluated. This study examines that design under group-aware evaluation, complete 14-class attribution, downstream-classifier sensitivity, end-to-end cascading, external validation, and reconstruction-error ablation.
Based on this gap, the study addresses the following research questions:
  • RQ1: Can task-specific autoencoder representations provide compact and informative features for IIoT intrusion detection?
  • RQ2: Does adding reconstruction error to the latent representation improve binary intrusion detection compared with using latent features alone?
  • RQ3: How effectively can the proposed task-specific autoencoder framework support multiclass attack attribution under imbalanced IIoT traffic, including the complete 14-class setting?
To address these questions, the framework uses separate 16D benign-trained and 24D malicious-trained autoencoder representations for binary detection and attack attribution, respectively. CatBoost is the main classifier, with XGBoost and LightGBM used for sensitivity analysis, alongside group-aware, cascade, WUSTL-IIoT-2021 [20], ablation, and CPU-based evaluations.

3. A Hybrid Autoencoder–CatBoost Approach for IIoT Intrusion Detection

The proposed framework combines task-specific autoencoder-based representation learning with downstream CatBoost classification for IIoT intrusion detection. Two complementary representation-learning branches are used. In the binary branch, an autoencoder is trained exclusively on benign traffic, and the resulting latent representation is combined with the sample-wise reconstruction error before CatBoost classification. In the multiclass branch, a separate autoencoder is trained exclusively on malicious traffic, and its latent representation is used by CatBoost for attack-type attribution. This separation reflects the distinct objectives of binary intrusion detection and multiclass attack classification considered in recent IDS evaluation studies [37].
The evaluation includes an initial reference setting and an extended robustness setting on ML-EdgeIIoT, derived from Edge-IIoTset [10]. The initial setting retains the original binary task and 12-class malicious-only benchmark, excluding MITM and Fingerprinting. The extended setting restores all 14 attacks and uses five group-aware folds that keep identical 42-feature vectors within the same fold; both branches are then combined into a 15-label cascade. The framework is also retrained on WUSTL-IIoT-2021 [20] for external validation, and CPU-based latency, throughput, and memory overhead are reported. The overall workflow of the proposed framework is illustrated in Figure 1.
Figure 1. Overview of the proposed task-specific IDS framework and evaluation protocol. Separate autoencoders learn benign-oriented and attack-oriented representations for binary intrusion detection and multiclass attack attribution, respectively. The extended evaluation includes group-aware splitting, a 15-label cascade, external validation, and computational assessment (Arrows indicate the direction of data flow and processing through the framework).

3.1. Datasets and Preprocessing

Experiments use ML-EdgeIIoT, a processed configuration derived from Edge-IIoTset [10], containing 157,800 samples and 42 numerical predictors. Attack_label defines binary benign-versus-attack detection, whereas Attack_type defines multiclass attack attribution. Table 2 summarizes the evaluation settings. Only the 42 numerical predictors were retained [38] and normalized to the [0, 1] range using Min-Max scaling (Equation (1)) [39]. For ML-EdgeIIoT, Attack_label and Attack_type are used only as target variables and are excluded from the predictor matrix; the retained 42 predictors are numerical, so no additional categorical-feature encoding is applied.
x scaled = x − min x max x − min x
where the scaling parameters are estimated on training data and applied unchanged to validation and test data.
Table 2. Dataset configurations and evaluation settings used in the study.
For the extended ML-EdgeIIoT evaluation, MITM and Fingerprinting are restored, yielding 14 attack classes. Among 157,800 samples, 152,029 unique 42-feature vectors were identified; 7149 samples belonged to duplicated-vector groups, and 188 groups contained conflicting Attack_type labels. Identical feature vectors are therefore kept within the same fold to prevent duplicate leakage between training and test data.
External validation uses WUSTL-IIoT-2021 [20], containing 1,194,464 samples: 1,107,448 normal and 87,016 attack observations. After removing StartTime, LastTime, SrcAddr, DstAddr, sIpId, and dIpId, 41 numerical predictors remain. Target defines binary Normal/Attack detection, while Traffic provides four attack classes: DoS, Reconn, CommInj, and Backdoor, with Normal retained for the five-label cascade. A group-safe approximately 60%/20%/20% train-validation-test split keeps identical 41-feature vectors within the same partition; preprocessing is fitted on training data only.

3.2. Autoencoder-Based Representation Learning

A fully connected autoencoder learns a compact representation of the numerical traffic features. For an input x ∈ ℝn, the encoder maps x to a latent vector z as defined in Equation (2).
z = fθ(x)
The decoder reconstructs the input according to Equation (3).
x ˆ = g ϕ z
Training minimizes the mean squared reconstruction error defined in Equation (4).
L AE x = 1 n ∑ i = 1 n x i − x ^ i 2
This objective encourages the latent representation to preserve informative structure while reducing redundancy [40,41].
For ML-EdgeIIoT, n = 42. The binary branch uses a 16-dimensional latent space selected during development using validation data only, with test data excluded from selection. The multiclass branch uses a 24-dimensional latent space retained as the strongest evaluated development setting. Both dimensions are kept fixed in the extended experiments to avoid test-driven re-tuning.
For WUSTL-IIoT-2021, the same 16- and 24-dimensional latent sizes are retained without re-tuning; only the input dimension changes to n = 41.

3.3. Binary Intrusion Detection Setting

The binary branch is evaluated under the initial reference benchmark and the extended five-fold group-aware protocol. Its autoencoder is trained only on benign traffic so that reconstruction error can reflect deviation from learned normal patterns [34]. Each sample is represented by the 16-dimensional latent vector and a scalar reconstruction error defined in Equation (5).
e(x) = LAE(x)
These are concatenated to form the final feature vector described by Equation (6).
x ~ = z ; e ( x )
The fused representation is classified with CatBoost, retained as the fixed downstream learner of the proposed pipeline. CatBoost uses ordered boosting to reduce prediction shift during training [42]. Binary training minimizes the weighted log-loss in Equation (7).
L bin = − 1 N ∑ i = 1 N w y i y i log p i + 1 − y i log 1 − p i
where N is the number of training samples, yi is the binary label, pi is the predicted attack probability, and w y i is the class weight associated with yi.
Binary performance is evaluated using ROC-AUC [43,44], which summarizes benign-versus-attack discrimination across decision thresholds, as defined in Equation (8).
AUC = ∫ 0 1 TPR ( FPR )   d ( FPR )
The true positive rate (TPR), also referred to as recall, and the false positive rate (FPR) are defined in Equations (9) and (10), respectively.
TPR = TP TP + FN
FPR = FP FP + TN
Average precision (AP) is also reported because it summarizes the precision–recall trade-off under class imbalance [44], as defined in Equation (11).
AP = ∑ n = 1 N ( R n − R n − 1 ) P n
where Pn and Rn denote the precision and recall values at the n-th operating point, respectively, and N is the number of evaluated points on the precision–recall curve.
For the extended evaluation, ML-EdgeIIoT is assessed using five group-aware outer folds, with identical 42-feature vectors kept within the same fold to prevent duplicate leakage. Each outer-training partition is further divided into group-safe inner training and validation subsets. The autoencoder is trained only on benign inner-training samples, with benign inner-validation samples used for early stopping. The resulting 16-dimensional latent features and reconstruction error are concatenated and provided to CatBoost.
CatBoost is trained on the inner-training subset only. The malicious-class weight is computed from the benign-to-malicious ratio in that subset, while the decision threshold is selected exclusively on inner validation by maximizing binary F1. The outer fold is used only for final evaluation. After all five folds, the out-of-fold predictions are concatenated to obtain global robustness metrics.

3.4. Multiclass Attack Classification Setting

The multiclass branch is evaluated in the initial 12-class reference and extended 14-class settings. Benign traffic is removed because this stage performs attack attribution; its dedicated autoencoder is therefore trained on malicious traffic to model differences among attack types. This task-specific choice is not assumed to be optimal. The initial setting excludes MITM and Fingerprinting because of their limited support, whereas the extended setting restores both classes, with 1214 and 1001 samples, respectively (Figure 2).
Figure 2. Distribution of the 14 malicious attack categories in ML-EdgeIIoT. MITM and Fingerprinting, highlighted by hatched bars, were excluded from the initial 12-class benchmark because of their limited support and reintroduced in the extended 14-class evaluation.
Both settings use the same malicious-only autoencoder with a 24-dimensional latent representation followed by CatBoost. For the extended evaluation, the established architecture and classifier settings are kept fixed. The same five outer group-aware folds used in the binary analysis are retained, keeping identical 42-feature vectors within the same fold. Predictions from malicious samples across all outer test folds are concatenated to obtain global out-of-fold performance.
For K attack classes, class probabilities are estimated using Equation (12), with K = 12 for the initial benchmark and K = 14 for the extended evaluation.
P y = k x = exp ( s k x ) ∑ j = 1 K exp ( s j x )
Training minimizes the multiclass cross-entropy loss defined in Equation (13).
L multi = − 1 m ∑ i = 1 m ∑ k = 1 K I y i = k log P y i = k ∣ x i
where I(yi = k) is an indicator function equal to 1 if sample i belongs to class k, and 0 otherwise.
Before multiclass filtering, ML-EdgeIIoT contains 157,800 samples: Normal (24,301), Backdoor (10,195), DDoS_HTTP (10,561), DDoS_ICMP (14,090), DDoS_TCP (10,247), DDoS_UDP (14,498), Fingerprinting (1001), MITM (1214), Password (9989), Port_Scanning (10,071), Ransomware (10,925), SQL_injection (10,311), Uploading (10,269), Vulnerability_scanner (10,076), and XSS (10,052). For the initial 12-class benchmark, Normal, MITM, and Fingerprinting are excluded, leaving 131,284 malicious samples; the counts of the 12 retained attack classes are unchanged. The extended 14-class benchmark restores MITM and Fingerprinting, yielding 133,499 malicious samples.

3.5. End-to-End Cascade and External Validation

The binary and multiclass branches are combined into an end-to-end cascade over mixed traffic. Samples predicted as Normal stop at the binary stage; samples predicted as Attack are forwarded to the 14-class attribution branch. The final ML-EdgeIIoT output therefore contains 15 labels: Normal plus 14 attacks. The same five outer group-aware folds are used, and binary-stage errors are retained in the final evaluation.
External validation is performed on WUSTL-IIoT-2021 [20] using the same task-specific AE-CatBoost design with 41 input features. The framework is retrained on WUSTL rather than transferred directly from ML-EdgeIIoT and is evaluated with the group-safe train-validation-test split defined in Section 3.1. Binary detection, four-class attack attribution, and the resulting five-label cascade are evaluated without WUSTL-specific hyperparameter re-tuning.

3.6. Evaluation Metrics

Multiclass performance is evaluated using accuracy, macro-F1, and one-vs-rest macro-AUC [45,46], as defined in Equations (14)–(19).
Accuracy = ∑ k = 1 K T P k N
For each class k:
Precision k = TP k TP k + FP k
Recall k = TP k TP k + FN k
F 1 k = 2 Precision k Recall k Precision k + Recall k
MacroF 1 = 1 K ∑ k = 1 K F 1 k
MacroAUC = 1 K ∑ k = 1 K A U C k
Macro-averaged metrics are emphasized because they provide a more informative view of IDS performance under class imbalance than accuracy alone [25,43].
For the extended 14-class and end-to-end evaluations, per-class precision, recall, F1-score, support, and Weighted-F1 are also reported. Cascade performance additionally includes attack recall, benign FPR, conditional attribution accuracy over true attacks forwarded to the multiclass stage, and strict end-to-end attack-type success, defined as the proportion of all true attacks that are both detected and correctly attributed.

3.7. Training, Reporting, and Computational Evaluation

All preprocessing and model selection use training and validation data only, with outer test folds reserved for final evaluation. The main extended ML-EdgeIIoT and WUSTL-IIoT-2021 experiments retain the established AE-CatBoost configurations without test-driven re-tuning [26]. A separate classifier-sensitivity analysis compares CatBoost, XGBoost, and LightGBM using identical AE features and the same outer folds. Each booster is evaluated over four depth/learning-rate combinations selected on inner validation only; the outer fold is evaluated once after selection.
For binary comparison, classifier selection uses validation ROC-AUC and the decision threshold is chosen by validation F1; for 14-class attribution, selection uses validation macro-F1. This analysis is reported separately and does not replace the fixed CatBoost results of the main framework. A separate binary ablation compares the 16D latent representation alone with the 17D latent-plus-reconstruction-error representation under the same five-fold group-aware protocol; within each fold, the same AE and CatBoost settings are used, so only the classifier input changes.
Deployment-oriented computational feasibility is assessed on a CPU-only system through single-sample inference latency, batched and full-cascade throughput, and process-memory overhead. Timing measurements are performed after model warm-up, while memory is measured in an isolated process using model-loading RSS increase and additional peak RSS during inference. These measurements characterize offline tabular inference rather than live packet-stream deployment.

4. Results

This section reports the initial reference benchmarks and the extended evaluations defined in Section 3. Results cover binary detection and reconstruction-error ablation, multiclass attack attribution, classifier sensitivity, the 15-label cascade, external validation on WUSTL-IIoT-2021, and offline CPU-based computational feasibility. For group-aware ML-EdgeIIoT experiments, global out-of-fold results are emphasized, with class-wise and fold-level summaries where relevant.

4.1. Binary Detection Results

The initial binary reference achieved ROC-AUC = 0.9776 and AP = 0.9960. Under the extended five-fold group-aware protocol, using the validation-selected threshold in each outer fold, global OOF performance reached ROC-AUC = 0.9925, AP = 0.9981, precision = 0.9799, F1 = 0.9829, attack recall = 0.9859, FNR = 0.0141, and benign FPR = 0.1113. Across folds, ROC-AUC averaged 0.9925 ± 0.0045 and AP 0.9981 ± 0.0018. Because the two settings use different evaluation protocols, these results are complementary rather than directly comparable.
Despite strong discrimination, the 11.13% benign FPR remains an important operational limitation. The corresponding ROC and Precision–Recall curves are shown in Figure 3.
Figure 3. Binary detection performance under the extended five-fold group-aware evaluation on ML-EdgeIIoT: (a) ROC curve, where the dashed diagonal line represents the no-skill (random-classifier) baseline; (b) Precision–Recall curve. Both curves are computed from global out-of-fold predictions, with each sample predicted only when assigned to a held-out outer fold.
These results show that the learned latent representation combined with reconstruction error retains strong binary discrimination under group-aware evaluation. However, the observed benign FPR confirms that further false-alarm reduction is still needed for operational use. A binary ablation under the same group-aware protocol compared the 16D latent representation alone with the 17D latent-plus-reconstruction-error representation. Adding reconstruction error increased global OOF F1 from 0.9800 to 0.9829 and ROC-AUC from 0.9912 to 0.9925, while reducing benign FPR from 0.1311 to 0.1113. These differences are descriptive; no statistical significance test was performed.

4.2. Multiclass Attack Classification Results

The initial 12-class reference benchmark achieved accuracy = 0.7226, macro-F1 = 0.6905, and macro-AUC = 0.9561. After restoring MITM and Fingerprinting, the extended 14-class group-aware evaluation achieved global OOF accuracy = 0.6654, macro-F1 = 0.6175, and macro-AUC = 0.9549. The relatively high macro-AUC indicates preserved score-level separability, whereas the lower hard-label metrics reflect the greater difficulty of complete attack attribution. No dedicated multiclass probability-calibration or class-specific threshold optimization was applied; therefore, the macro-AUC–macro-F1 gap is interpreted as a difference between score-level separability and hard-label classification performance rather than as evidence of well-calibrated class probabilities. Because the class composition and evaluation protocols differ, the two settings are complementary rather than directly comparable.
Class-wise performance remains heterogeneous. DDoS_ICMP and DDoS_UDP are among the most reliably attributed classes, whereas Fingerprinting and MITM remain more difficult, with F1-scores of 0.5966 and 0.4737, respectively. The row-normalized confusion matrix for the extended 14-class evaluation is shown in Figure 4.
Figure 4. Row-normalized confusion matrix for the extended 14-class five-fold group-aware evaluation on ML-EdgeIIoT. The matrix is computed from global out-of-fold predictions, with each sample evaluated only in its held-out outer fold. Values are percentages within each true class.
MITM remains particularly difficult to attribute, consistent with conflicting Attack_type labels observed among some identical feature vectors.
To assess sensitivity to the downstream booster, CatBoost, XGBoost, and LightGBM were compared under the controlled protocol defined in Section 3.7. For binary detection, global OOF performance was very similar across classifiers: ROC-AUC = 0.9927, 0.9927, and 0.9929, and F1 = 0.9820, 0.9821, and 0.9818, respectively. LightGBM yielded the lowest benign FPR (0.1047). Differences were larger for 14-class attribution, where macro-F1 reached 0.6634, 0.6819, and 0.6946, with macro-AUC = 0.9584, 0.9586, and 0.9604, respectively. Thus, binary discrimination shows little numerical variation across boosters, whereas attack attribution is more sensitive to the downstream classifier. These results are descriptive and do not imply statistical superiority of any classifier.
Because this sensitivity analysis uses an equal-budget inner-validation selection procedure, its CatBoost values are not intended to replace the fixed-configuration results reported in Table 3 and Table 4.
Table 3. Binary detection performance of the main AE-CatBoost framework under the initial and extended evaluation protocols.
Table 4. Multiclass attack-attribution performance of the main AE-CatBoost framework under the initial and extended evaluation protocols.
The controlled classifier-sensitivity results are summarized in Table 5.
Table 5. Controlled downstream-classifier sensitivity analysis under the five-fold group-aware protocol.

4.3. End-to-End Cascade Results

The complete 15-label cascade was evaluated on the full ML-EdgeIIoT dataset using the same five-fold group-aware protocol. Global out-of-fold results reached an accuracy of 0.6939, a macro-F1 of 0.6273, and a Weighted-F1 of 0.6793. The binary stage achieved an attack recall of 0.9859 with a benign FPR of 0.1113, while conditional attribution accuracy reached 0.6679. Overall, 65.84% of all true attacks were both detected and correctly attributed. Of 133,499 true attacks, 1885 were missed at the binary stage and 43,714 were misclassified after detection, indicating that most remaining errors arise during fine-grained attribution. Fingerprinting and MITM remained particularly difficult, with final recalls of 0.4426 and 0.3114, respectively. The corresponding end-to-end cascade results are summarized in Table 6.
Table 6. End-to-end 15-label cascade performance on ML-EdgeIIoT.
Conditional attribution accuracy is computed over true attacks forwarded to the multiclass stage, whereas strict end-to-end success requires both correct attack detection and correct attack-type attribution.

4.4. External Cross-Benchmark Validation on WUSTL-IIoT-2021

The AE-CatBoost framework was retrained and evaluated on WUSTL-IIoT-2021 using its 41-feature space and the group-safe split defined in Section 3.1, without dataset-specific hyperparameter re-tuning. Binary detection achieved ROC-AUC = 0.999995, AP = 0.999934, and F1 = 0.9990. Four-class attribution reached macro-F1 = 0.9767, while the five-label cascade achieved macro-F1 = 0.9408 and strict end-to-end success = 0.9993. These results show strong performance on the evaluated WUSTL setting but should be interpreted in light of its high-class separability and limited support for some rare classes. The corresponding cross-benchmark validation results are summarized in Table 7.
Table 7. External cross-benchmark performance on WUSTL-IIoT-2021.

4.5. Deployment-Oriented Computational Feasibility

Computational feasibility of the complete WUSTL-IIoT-2021 five-label cascade was evaluated on a CPU-only system with 8 physical cores, 16 logical cores, and 15.73 GB RAM. Median single-sample latency was 6.51 ms for true-benign and 9.35 ms for true-attack samples, with 95th-percentile latencies of 9.24 ms and 15.90 ms, respectively. Batched full-cascade inference reached a median throughput of 147,528 samples/s and an amortized time of 0.0068 ms/sample. Model loading increased process RSS by 22.20 MB, with an additional inference peak of 9.07 MB. These measurements describe offline tabular inference and exclude packet capture, feature extraction, and network I/O.
Model-storage size and training time were not measured, and no edge-device deployment was performed; these aspects remain outside the present computational evaluation. The corresponding deployment-oriented computational measurements are summarized in Table 8.
Table 8. Deployment-oriented computational measurements of the complete WUSTL-IIoT-2021 five-label cascade.

5. Discussion

This study evaluates a modular IDS based on task-specific autoencoder representation learning and downstream boosting. Separate autoencoders are trained on benign traffic for binary detection and malicious traffic for attack attribution.
The extended evaluation examines this design under group-aware splitting, all 14 attack classes, reconstruction-error ablation, end-to-end cascading, classifier sensitivity, external validation, and CPU-based computational analysis.
For ML-EdgeIIoT, the extended binary evaluation achieved ROC-AUC = 0.9925 and AP = 0.9981. The reconstruction-error ablation showed modest gains over latent-only features, increasing global OOF F1 from 0.9800 to 0.9829 and reducing the benign FPR from 0.1311 to 0.1113, although this remaining false-positive rate is still an operational limitation. The complete 14-class evaluation was more difficult, with macro-F1 = 0.6175 and macro-AUC = 0.9549; MITM and Fingerprinting remained among the hardest classes, consistent with conflicting labels observed among some identical feature vectors. The classifier-sensitivity analysis showed little variation in binary ROC-AUC across CatBoost, XGBoost, and LightGBM (0.9927–0.9929), but larger differences in 14-class macro-F1 (0.6634–0.6946), indicating greater sensitivity of attack attribution to the downstream booster.
The 15-label cascade achieved attack recall = 0.9859, conditional attribution accuracy = 0.6679, and strict end-to-end success = 0.6584, indicating that most residual errors occurred during attack attribution. On WUSTL-IIoT-2021, retraining without WUSTL-specific re-tuning yielded ROC-AUC = 0.999995 for binary detection, macro-F1 = 0.9767 for four-class attribution, and macro-F1 = 0.9408 with strict success = 0.9993 for the five-label cascade. These results reflect the evaluated WUSTL setting, which shows greater class separability and limited support for some rare classes, and should not be directly compared with ML-EdgeIIoT. CPU-only inference showed median latencies of 6.51 ms for true-benign and 9.35 ms for true-attack samples, with a batched throughput of 147,528 samples/s; these measurements exclude packet capture, feature extraction, and network I/O.
To contextualize the main ML-EdgeIIoT evaluation, Table 9 summarizes representative studies using Edge-IIoTset or closely related hybrid IDS settings. The WUSTL external validation, classifier-sensitivity analysis, and end-to-end cascade address complementary evaluation questions and are therefore discussed separately rather than treated as directly comparable benchmarks.
Table 9. Methodological positioning of the proposed framework against representative IDS studies.
Table 9 positions the proposed framework against representative IDS studies without implying direct numerical comparability across heterogeneous protocols. Among the AE-based studies considered, the main distinguishing design choice is the use of separate autoencoders trained on benign and malicious traffic for binary detection and attack attribution, respectively. The contribution therefore lies in this task-specific modular design and its evaluation under group-aware splitting, complete 14-class attribution, reconstruction-error ablation, classifier sensitivity, end-to-end cascading, external validation, and computational analysis, rather than in numerical superiority over prior work. The findings should be interpreted within several limitations. ML-EdgeIIoT is a processed Edge-IIoTset configuration and contains duplicated feature vectors, including some with conflicting attack labels, which may complicate fine-grained attribution. External validation was limited to WUSTL-IIoT-2021, where the observed strong class separability and limited support for some rare attacks restrict broader generalization claims. The computational evaluation covers offline tabular inference only and excludes packet capture, feature extraction, and network I/O; it therefore does not establish real-time deployment. Robustness to adversarial perturbations, noisy or missing features, concept drift, and previously unseen attack types was not evaluated and remains an important direction for future work.
Taken together, the results show strong binary discrimination but greater difficulty in fine-grained attack attribution. The 14-class and cascade evaluations highlight persistent challenges associated with rare classes and label ambiguity, while the binary FPR indicates that false-alarm reduction remains necessary. These findings support the framework as a modular IDS design while leaving room for further validation and operational refinement.

6. Conclusions and Future Work

This study developed a modular IIoT intrusion-detection framework based on task-specific autoencoder representation learning and CatBoost classification. Separate autoencoders are trained on benign traffic for binary detection and malicious traffic for attack attribution. The framework was evaluated through initial reference benchmarks and complementary group-aware, 14-class, cascade, classifier-sensitivity, external, ablation, and computational analyses.
The research questions can be answered as follows. For RQ1, the task-specific autoencoders reduced the 42-feature input to 16D and 24D representations that supported binary detection and attack attribution. For RQ2, the ablation analysis showed that adding reconstruction error to the 16D latent representation produced modest overall gains: global OOF F1 increased from 0.9800 to 0.9829, while the benign FPR decreased from 0.1311 to 0.1113. For RQ3, multiclass attack attribution remained challenging under class imbalance. The initial 12-class reference achieved macro-F1 = 0.6905 and macro-AUC = 0.9561, while the complementary 14-class group-aware evaluation achieved macro-F1 = 0.6175 and macro-AUC = 0.9549, indicating the greater difficulty of complete fine-grained attribution.
The complementary analyses provide further context. Binary performance varied little across CatBoost, XGBoost, and LightGBM, whereas 14-class attribution was more sensitive to the downstream classifier. The 15-label cascade achieved strict end-to-end attack-type success = 0.6584, while WUSTL-IIoT-2021 validation yielded five-label macro-F1 = 0.9408 and strict success = 0.9993; these external results should be interpreted in light of the greater class separability and limited support of some rare WUSTL classes. CPU measurements also demonstrated low-cost offline tabular inference, but do not establish real-time packet-stream deployment. Taken together, these findings support the framework as a transparent and extensible IDS design rather than as a universally optimal solution.
Future work should extend validation to additional IoT/IIoT datasets and investigate strategies for reducing false alarms and improving rare-class attribution. Live packet-stream evaluation should include packet capture, feature extraction, and network I/O to assess operational latency beyond offline tabular inference. Further work should examine adversarial robustness and integrate explainable AI methods such as SHAP or LIME to support feature- and sample-level interpretation.
A complementary direction will specifically assess the framework under zero-day conditions using open-set and leave-one-attack-type-out protocols, with explicit evaluation of unknown-attack detection and rejection.

Author Contributions

S.F. designed the proposed approach, conducted the experiments, and prepared the initial draft of the manuscript. D.B. contributed to the conceptual discussion, refined the methodology, and critically revised the manuscript. F.T. and Y.L. reviewed the paper and provided valuable feedback for improvement. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The data analyzed in this study were derived from two publicly available datasets. The primary experiments used the processed file ML-EdgeIIoT-dataset.csv, derived from Edge-IIoTset. The original Edge-IIoTset dataset is publicly available at https://www.kaggle.com/datasets/mohamedamineferrag/edgeiiotset-cyber-security-dataset-of-iot-iiot (accessed on 21 September 2026) and is described in Ferrag et al. [10]. https://doi.org/10.1109/ACCESS.2022.3165809. External validation used the publicly available WUSTL-IIoT-2021 dataset, identified by DOI: https://doi.org/10.21227/yftq-n229. The preprocessing and evaluation code used in this study, together with the saved split manifests are available from the corresponding author upon reasonable request.

Acknowledgments

During the preparation of this manuscript, the authors used ChatGPT [5.3] for writing enhancement and to assist in preparing the schematic framework illustration in Figure 1. Figure 2, Figure 3 and Figure 4 are data-driven visualizations generated from the authors’ datasets, experimental predictions, and analysis code; ChatGPT was not used to generate experimental results or performance values. The authors reviewed and edited all AI-assisted content and take full responsibility for the publication.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Zhukabayeva, T.; Zholshiyeva, L.; Karabayev, N.; Khan, S.; Alnazzawi, N. Cybersecurity Solutions for Industrial Internet of Things–Edge Computing Integration: Challenges, Threats, and Future Directions. Sensors 2025, 25, 213. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Knapp, E.D.; Langill, J.T. Industrial Network Security: Securing Critical Infrastructure Networks for Smart Grid, SCADA, and Other Industrial Control Systems, 2nd ed.; Syngress: Waltham, MA, USA, 2015. [Google Scholar]
  3. Ismail, S.; Dandan, S.; Qushou, A. Intrusion Detection in IoT and IIoT: Comparing Lightweight Machine Learning Techniques Using TON_IoT, WUSTL-IIOT-2021, and Edge-IIoTset Datasets. IEEE Access 2025, 13, 73468–73485. [Google Scholar] [CrossRef] [Scilit]
  4. Al-Garadi, M.A.; Mohamed, A.; Al-Ali, A.K.; Du, X.; Ali, I.; Guizani, M. A Survey of Machine and Deep Learning Methods for Internet of Things Security. IEEE Commun. Surv. Tutor. 2020, 22, 1646–1685. [Google Scholar] [CrossRef] [Scilit]
  5. Bansal, K.; Singhrova, A. Review on Intrusion Detection System for IoT/IIoT: A Brief Study. Multimed. Tools Appl. 2024. [Google Scholar] [CrossRef] [Scilit]
  6. Rehman, H.M.R.U.; Liaquat, S.; Gul, M.J.; Jhandir, M.Z.; Gavilanes, D.; Vergara, M.M.; Ashraf, I. A Systematic Literature Study of Machine Learning Techniques-Based Intrusion Detection: Datasets, Models, Challenges, and Future Directions. J. Big Data 2025, 12, 264. [Google Scholar] [CrossRef] [Scilit]
  7. Pinto Neto, E.C.; Dadkhah, S.; Ferreira, R.; Zohourian, A.; Lu, R.; Ghorbani, A.A. CICIoT2023: A Real-Time Dataset and Benchmark for Large-Scale Attacks in IoT Environments. Sensors 2023, 23, 5941. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Rahman, M.M.; Al Shakil, S.; Mustakim, M.R. A Survey on Intrusion Detection Systems in IoT Networks. Cyber Secur. Appl. 2025, 3, 100082. [Google Scholar] [CrossRef] [Scilit]
  9. Li, J.; Othman, M.S.; Chen, H.; Yusuf, L.M. Optimizing IoT Intrusion Detection Systems: Feature Selection versus Feature Extraction in Machine Learning. J. Big Data 2024, 11, 24. [Google Scholar] [CrossRef] [Scilit]
  10. Ferrag, M.A.; Friha, O.; Hamouda, D.; Maglaras, L.; Janicke, H. Edge-IIoTset: A New Comprehensive Realistic Cyber Security Dataset of IoT and IIoT Applications for Centralized and Federated Learning. IEEE Access 2022, 10, 40281–40306. [Google Scholar] [CrossRef] [Scilit]
  11. Pan, L.; Liu, N.; Zheng, C. Addressing Class Imbalance in Intrusion Detection: A Comprehensive Evaluation of Machine Learning Approaches. Electronics 2025, 14, 69. [Google Scholar] [CrossRef] [Scilit]
  12. García, J.; Entrena, J.; Alesanco, Á. Empirical Evaluation of Feature Selection Methods for Machine Learning-Based Intrusion Detection in IoT Scenarios. Internet Things 2024, 28, 101367. [Google Scholar] [CrossRef] [Scilit]
  13. Ruiz-Villafranca, S.; Roldán-Gómez, J.; Carrillo-Mondéjar, J.; Martínez, J.L.; Gañán, C.H. WFE-Tab: Overcoming Limitations of TabPFN in IIoT-MEC Environments with a Weighted Fusion Ensemble-TabPFN Model for Improved IDS Performance. Future Gener. Comput. Syst. 2025, 166, 107707. [Google Scholar] [CrossRef] [Scilit]
  14. Alaghbari, K.A.; Lim, H.-S.; Saad, M.H.M.; Yong, Y.S. Deep Autoencoder-Based Integrated Model for Anomaly Detection and Efficient Feature Extraction in IoT Networks. IoT 2023, 4, 345–365. [Google Scholar] [CrossRef] [Scilit]
  15. Moucharraf, M.; Ridouani, M.; Salahdine, F.; Kaabouch, N. AutoBoost-IoT: A Hybrid Model for Intrusion Detection in IoT Networks. Future Internet 2026, 18, 229. [Google Scholar] [CrossRef] [Scilit]
  16. Hasan, T.; Hossain, A.; Ansari, M.Q.; Syed, T.H. Enhanced Intrusion Detection in IIoT Networks: A Lightweight Approach with Autoencoder-Based Feature Learning. In Proceedings of the 10th International Conference on Internet of Things, Big Data and Security (IoTBDS 2025); SciTePress: Porto, Portugal, 2025; pp. 207–214. [Google Scholar] [CrossRef] [Scilit]
  17. Shwaysh, M.M.; Hussain, A.-S.T.; Salih, S.Q.; Almulaisi, T.A.; Radhi, A.D.; Majdi, H.S.; Desa, H. Adaptive Hybrid Information Gain and Autoencoder-Based Feature Selection with Ensemble Recurrent Extreme Learning Machine for Enhanced Network Intrusion Detection Systems. J. Netw. Syst. Manag. 2026, 34, 1. [Google Scholar] [CrossRef] [Scilit]
  18. Kuraś, P.; Bolanowski, M.; Łoza, M. Optimization of Machine Learning Models for Effective Anomaly Detection in Industrial IoT Systems. Adv. Sci. Technol. Res. J. 2026, 20, 203–221. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Abinayaa, S.S.; Arumugam, P.; Mohan, D.B.; Rajendran, A.; Lashab, A.; Wei, B.; Guerrero, J.M. Securing the Edge: CatBoost Classifier Optimized by the Lyrebird Algorithm to Detect Denial of Service Attacks in Internet of Things-Based Wireless Sensor Networks. Future Internet 2024, 16, 381. [Google Scholar] [CrossRef] [Scilit]
  20. Zolanvari, M. WUSTL-IIoT-2021 Dataset for IIoT Cybersecurity Research. IEEE DataPort 2021. [Google Scholar] [CrossRef]
  21. Ba, A.; Add, M. Machine Learning for Intrusion Detection in IIoT: A Comprehensive Review. Procedia Comput. Sci. 2025, 272, 100–107. [Google Scholar] [CrossRef] [Scilit]
  22. Qaddos, A.; Yaseen, M.U.; Al-Shamayleh, A.S.; Imran, M.; Akhunzada, A.; Alharthi, S.Z. A Novel Intrusion Detection Framework for Optimizing IoT Security. Sci. Rep. 2024, 14, 21789. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Su, Y.-K.; Tseng, C.-J. A Physics-Informed Benchmarking Framework for Machine Learning and Tree-Based Ensembles in IIoT-Enabled Predictive Maintenance. Sensors 2026, 26, 5026. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Liu, C.-L.; Weng, P.-H.; Tseng, C.-J. Harnessing Heterogeneous Graph Neural Networks for Dynamic Job-Shop Scheduling Problem Solutions. Comput. Ind. Eng. 2025, 203, 111060. [Google Scholar] [CrossRef] [Scilit]
  25. Jamshidi, S.; Nafi, K.W.; Nikanjam, A.; Khomh, F. Evaluating Machine Learning-Driven Intrusion Detection Systems in IoT: Performance and Energy Consumption. Comput. Ind. Eng. 2025, 204, 111103. [Google Scholar] [CrossRef] [Scilit]
  26. Shaikhanova, A.; Kuznetsov, O.; Tokkuliyeva, A.; Ayapbergenov, K.; Olzhas, S.; Danir, T. Security Audit of IoT Device Networks: A Reproducible Machine Learning Framework for Threat Detection and Performance Benchmarking. Sensors 2025, 25, 7519. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Fraihat, S.; Yaseen, Q.; Sanjalawe, Y.; Abu-Errub, A.; Makhadmeh, S.N.; Al-Betar, M.A. Intrusion Detection in Industrial Internet of Things Network Using Feature Optimization and Hybrid Deep Learning. Discov. Internet Things 2026, 6, 34. [Google Scholar] [CrossRef] [Scilit]
  28. Yang, K.; Wang, J.; Li, M. An Improved Intrusion Detection Method for IIoT Using Attention Mechanisms, BiGRU, and Inception-CNN. Sci. Rep. 2024, 14, 19339. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Ishtiaq, W.; Zannat, A.; Parvez, A.H.M.S.; Hossain, M.A.; Kanchan, M.H.; Tarek, M.M. CST-AFNet: A Dual Attention-Based Deep Learning Framework for Intrusion Detection in IoT Networks. Array 2025, 26, 100501. [Google Scholar] [CrossRef] [Scilit]
  30. Ruiz-Villafranca, S.; Roldán-Gómez, J.; Castelo Gómez, J.M.; Carrillo-Mondéjar, J.; Martínez, J.L. A TabPFN-Based Intrusion Detection System for the Industrial Internet of Things. J. Supercomput. 2024, 80, 20080–20117. [Google Scholar] [CrossRef] [Scilit]
  31. Du, Y.; Ning, A.; Cheng, P.; Kumar, R.; Du, X. Industrial Internet of Things Intrusion Detection Based on a Hybrid Model of Pearson–Deep Neural Network and Transformer. Eng. Appl. Artif. Intell. 2026, 171, 114304. [Google Scholar] [CrossRef] [Scilit]
  32. Khan, M.Z.; Reshi, A.A.; Shafi, S.; Aljubayri, I. An Adaptive Hybrid Framework for IIoT Intrusion Detection Using Neural Networks and Feature Optimization Using Genetic Algorithms. Discov. Sustain. 2025, 6, 382. [Google Scholar] [CrossRef] [Scilit]
  33. Aldaej, A.; Ullah, I.; Ahanger, T.A.; Atiquzzaman, M. Ensemble Technique of Intrusion Detection for IoT-Edge Platform. Sci. Rep. 2024, 14, 11703. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Narmadha, S.; Balaji, N.V. Improved Network Anomaly Detection System Using Optimized Autoencoder–LSTM. Expert Syst. Appl. 2025, 273, 126854. [Google Scholar] [CrossRef] [Scilit]
  35. Dong, S.; Xia, Y.; Wang, T. Network Abnormal Traffic Detection Framework Based on Deep Reinforcement Learning. IEEE Wirel. Commun. 2024, 31, 185–193. [Google Scholar] [CrossRef] [Scilit]
  36. Naha, R.; Kamruzzaman, J.; Karmakar, G.; Teng, S.W.; Shah, R.; Bakaul, M.; Dong, S. IoT and ML for Distributed Renewable Energy Storage and Generation Management: A Review. ACM Comput. Surv. 2026, 58, 1–45. [Google Scholar] [CrossRef] [Scilit]
  37. Alharthi, A.; Alaryani, M.; Kaddoura, S. A Comparative Study of Machine Learning and Deep Learning Models in Binary and Multiclass Classification for Intrusion Detection Systems. Array 2025, 26, 100406. [Google Scholar] [CrossRef] [Scilit]
  38. Talukder, M.A.; Islam, M.M.; Uddin, M.A.; Hasan, K.F.; Sharmin, S.; Alyami, S.A.; Moni, M.A. Machine Learning-Based Network Intrusion Detection for Big and Imbalanced Data Using Oversampling, Stacking Feature Embedding and Feature Extraction. J. Big Data 2024, 11, 33. [Google Scholar] [CrossRef] [Scilit]
  39. Ruhland, J.B.; Masoudian, I.; Heider, D. Enhancing Deep Neural Network Training through Learnable Adaptive Normalization. Knowl.-Based Syst. 2025, 326, 113968. [Google Scholar] [CrossRef] [Scilit]
  40. Berahmand, K.; Daneshfar, F.; Salehi, E.S.; Li, Y.; Xu, Y. Autoencoders and Their Applications in Machine Learning: A Survey. Artif. Intell. Rev. 2024, 57, 28. [Google Scholar] [CrossRef] [Scilit]
  41. Laakom, F.; Raitoharju, J.; Iosifidis, A.; Gabbouj, M. Reducing Redundancy in the Bottleneck Representation of Autoencoders. Pattern Recognit. Lett. 2024, 178, 202–208. [Google Scholar] [CrossRef] [Scilit]
  42. Prokhorenkova, L.; Gusev, G.; Vorobev, A.; Dorogush, A.V.; Gulin, A. CatBoost: Unbiased Boosting with Categorical Features. In Advances in Neural Information Processing Systems 31 (NeurIPS 2018); Curran Associates, Inc.: Red Hook, NY, USA, 2018; pp. 6638–6648. [Google Scholar]
  43. Tian, J.; Zhu, H. Evaluating the Efficacy of AI-Driven Intrusion Detection Systems in IoT: A Review of Performance Metrics and Cybersecurity Threats. PeerJ Comput. Sci. 2025, 11, e3352. [Google Scholar] [CrossRef] [Scilit]
  44. Saito, T.; Rehmsmeier, M. The Precision–Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets. PLoS ONE 2015, 10, e0118432. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  45. Farhadpour, S.; Warner, T.A.; Maxwell, A.E. Selecting and Interpreting Multiclass Loss and Accuracy Assessment Metrics for Classifications with Class Imbalance: Guidance and Best Practices. Remote Sens. 2024, 16, 533. [Google Scholar] [CrossRef] [Scilit]
  46. Harbecke, D.; Chen, Y.; Hennig, L.; Alt, C. Why Only Micro-F1? Class Weighting of Measures for Relation Classification. In Proceedings of NLP Power! The First Workshop on Efficient Benchmarking in NLP; Association for Computational Linguistics: Dublin, Ireland, 2022; pp. 32–41. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Article metric data becomes available approximately 24 hours after publication online.