Skip to Content
MathematicsMathematics
  • Article
  • Open Access

30 September 2026

23 Pages

Flow Exporter Provenance as a Major Confounder in Cross-Dataset IoT Intrusion Detection

and
Department of Information Technology, College of Computer, Qassim University, Buraydah 51452, Saudi Arabia
*
Author to whom correspondence should be addressed.

Abstract

Machine-learning intrusion detection for the Internet of Things (IoT) routinely exceeds 99% accuracy on single datasets but fails when transferred to new networks or flow exporters. We formalize five failure modes, define a 23-feature canonical schema, and adapt three datasets (CICIoT2023, TON_IoT, Bot-IoT) to build a full six-pair transfer matrix across three classifiers, each with three random seeds. Random forest attains the highest mean AUC (0.740), yet its balanced accuracy collapses to chance (0.50–0.51) exclusively on CICFlowMeter-sourced pairs while remaining strong (0.72–0.87) on Zeek- and Argus-sourced pairs. This exporter-dependent collapse is reproduced across structurally unrelated classifiers, establishing it as our central finding. A controlled synthetic test confirms that CICFlowMeter’s directional heuristic destroys variance (up to 104 reduction) and causes a small, consistent transfer cost (+0.004 AUC), though full degradation requires compounded effects. Aggregate distributional distance does not predict transfer success (r = −0.17), ruling out a simple divergence explanation. Prior correction provides the largest ablation gain. Source-domain coverage is necessary but not sufficient for transfer, and we observed no universal sample-count threshold. Flow exporter provenance emerges as a major upstream confounder and a directly implicated contributing mechanism, though exporter identity is confounded with dataset identity and is not established as the sole determinant of cross-dataset performance.

1. Introduction

The Internet of Things (IoT) now underpins infrastructure across manufacturing, healthcare, transport, and utilities, deploying a large number of resource-constrained endpoints on a shared network infrastructure [1,2]. The security consequences are substantial, and network-based intrusion detection has consequently attracted sustained research attention.
A considerable body of work applies machine learning to this problem, with reported detection accuracies frequently exceeding 99% on established benchmark datasets [3,4,5]. In most cases, these figures come from within-dataset evaluation, where training and test partitions are drawn from a single network capture. Such protocols measure a model’s capacity to generalize across draws from one distribution; they do not measure generalization across distributions. Independent evaluations that apply trained models to held-out datasets report F1 degradations of 30 to 50 percentage points [6,7,8], indicating that within-dataset accuracy weakly predicts operational performance.
The mechanisms responsible for this degradation remain incompletely characterized. Prior work has documented the magnitude of the gap [8] and identified dataset-level biases that contribute to it [7] but has not separated the individual mechanisms, established their relative contributions, or directly tested proposed causal explanations. The present study addresses this gap through a combination of a complete pairwise transfer matrix across three real datasets and three classifier families, and a controlled synthetic experiment that isolates one proposed mechanism and measures its causal contribution directly.

1.1. Motivating Observation

A GRU classifier trained on CICIoT2023 attained F1 = 0.929 on a held-out partition of the same dataset. Applied to Bot-IoT, a dataset collected by an adjacent research community in the same period and targeting comparable device classes, the same model attained an AUC of 0.562, not appreciably different from random assignment. An identically configured model trained on TON_IoT attained an AUC of 0.935 on Bot-IoT. Because the architecture, hyperparameters, and canonical feature set were held constant, the difference is attributable to the source dataset. This pattern is not specific to the GRU: a random forest trained on TON_IoT or Bot-IoT achieves AUC above 0.92 on several transfer pairs, while the identical model trained on CICIoT2023 collapses to near-chance balanced accuracy. CICIoT2023 was generated using CICFlowMeter, which does not export true bidirectional byte counts; TON_IoT and Bot-IoT were generated using Zeek and Argus, respectively, both of which do.

1.2. Contributions

The contributions of this work are as follows:
(1)
A formal taxonomy of five cross-dataset failure modes, each with a measurable diagnostic criterion.
(2)
A 23-feature canonical flow schema recoverable from any standard flow exporter, with adapters for CICIoT2023, TON_IoT, and Bot-IoT that exclude network identifiers by construction, and a fully disclosed label-to-family mapping.
(3)
A pairwise transfer matrix over six directed dataset pairs and three classifier families (logistic regression, random forest, GRU), with three random seeds per configuration. We report five of the six pairs with paired seed-level significance tests; we report the CIC→TON pair descriptively because its original run used an intermediate configuration that we cannot recompute without the raw data.
(4)
Evidence that flow-exporter provenance is strongly associated with transfer failure across structurally unrelated model families, together with an explicit analysis of why this association cannot be interpreted as fully isolated causation in the present design.
(5)
A direct, controlled test of the proposed causal mechanism, CICFlowMeter’s SYN/ACK directional-feature heuristic, using synthetic data with known ground truth, together with an empirical test of the divergence term in a standard domain-adaptation risk.
(6)
A label-budget analysis showing that, on pairs where the source threshold is miscalibrated, 50 labeled target-domain flows recover 98.7% of the achievable recalibration gain, while on pairs where the source threshold is already well calibrated, the procedure does not improve over zero-shot performance.
The rest of this paper is organized as follows: Section 2 reviews related work. Section 3 presents the theoretical framework, datasets, and experimental design, including the synthetic mechanism test. Section 4 reports results. Section 5 discusses the interpretation and limitations, and Section 6 concludes.

3. Materials and Methods

3.1. Taxonomy of Cross-Dataset Failure Modes

Let S = (X, Y, PS) denote a source domain with canonical feature space X, a subset of R23, binary label space Y = {0, 1}, and joint distribution PS, and let T = (X, Y, PT) denote a target domain. A classifier h: X maps to [0, 1] and is trained on labeled samples from S and evaluated on samples from T. Five mechanisms may cause the target risk e_T (h) to exceed the source risk e_S (h).
Definition 1 (capture-artifact leakage).
A feature f constitutes a capture artifact if its mutual information with Y given the source capture configuration CS is zero, while its mutual information with CS alone is positive. Retaining network identifiers produces within-dataset AUC values approaching unity; removing them reduced cross-dataset AUC from 1.000 to 0.300 in preliminary experiments.
Definition 2 (covariate shift).
PS(X) differs from PT(X) while P(Y given X) is invariant. In network measurement, this arises from differences in link capacity, maximum transmission unit, and exporter aggregation policy.
Definition 3 (label-prior shift).
PS(Y) differs from PT(Y) while P(X given Y) is invariant. The attack priors of the three datasets used here are 0.978 (CICIoT2023, prior to correction), 0.763 (TON_IoT), and 0.9983 (Bot-IoT, within the 1.2 million-flow processed subsample).
Definition 4 (concept shift).
PS(Y given X) differs from PT(Y given X). The same feature vector carries different label semantics across domains.
Definition 5 (threshold miscalibration).
For a model whose score function ranks target-domain samples correctly, a threshold calibrated on S may nonetheless be suboptimal on T. Distinguished from representation failure by high target AUC combined with low F1 at the source-selected threshold.

3.2. Diagnostic Procedure

Step 1. Train with identifiers retained. A within-dataset AUC approaching unity indicates that identifiers are acting as predictors.
Step 2. Compute the maximum mean discrepancy (MMD) [30] between source and target feature distributions after per-domain normalization. Section Testing the Divergence Term Empirically reports that this diagnostic, while theoretically motivated, does not predict transfer success in practice on our data.
Step 3. Compare class priors. A difference exceeding 0.15 warrants correcting the training set priors.
Step 4. Compare target AUC with F1 at the source threshold. High AUC combined with low F1 indicates threshold miscalibration rather than representation failure.
Step 5. Decompose target performance by attack family. Low target F1 together with adequate source coverage points toward concept shift; low target F1 together with sparse source coverage points toward insufficient coverage. The two are not separable by sample count alone, and family-level transfer depends jointly on coverage, target support, family separability, prior shift, and threshold calibration.

3.3. Canonical Feature Schema

The canonical schema was constructed to satisfy three requirements: recoverability from any standard flow exporter; exclusion of capture artifacts; and dimensional consistency across exporters. Table 1 summarizes the resulting 23 features.
Table 1. Canonical flow feature schema. Features marked as approximated in CICIoT2023 are estimated from CICFlowMeter fields rather than measured directly.
Table 1 was constructed under three explicit criteria: (i) recoverability from CICFlowMeter, Zeek, and Argus; (ii) exclusion of capture artifacts as defined above; and (iii) dimensional consistency across exporters. Features considered but excluded include raw TCP flag counts other than SYN/ACK, TCP window size, inter-arrival-time moments beyond duration, exporter-specific record identifiers, IP/port fields, timestamps, MAC/VLAN metadata, and payload-derived statistics. We excluded these either because at least one exporter does not provide them in a comparable form because they are capture artifacts or because they require target-domain labels or packet payloads to compute. The resulting 23-feature schema is therefore a compromise between expressiveness and cross-exporter portability, not an exhaustive representation of flow-level information.
CICIoT2023 adapter. Flow duration is not exported and was estimated as the product of inter-arrival time and packet count. We estimated the forward fraction as the SYN count divided by the sum of SYN and ACK counts, clipped to the range 0.3 to 0.7. Of the 23 features, 8 are therefore approximated rather than measured.
TON_IoT adapter. Zeek connection logs provide src_bytes, dst_bytes, and duration directly, so we can compute all 23 features without approximation.
Bot-IoT adapter. Argus provides pkts, bytes, dur, srate, and drate. We estimated the forward fraction as srate divided by the sum of srate and drate, clipped to 0.3 to 0.7. We blocked the fields pkSeqID, seq, stime, ltime, saddr, and daddr as identifiers.
We excluded capture artifacts before feature construction, following Definition 1. In Bot-IoT, we explicitly blocked pkSeqID, seq, stime, ltime, saddr, and daddr. pkSeqID and seq are row/record identifiers that encode file order and capture segmentation; stime and ltime are absolute timestamps tied to the testbed clock; saddr and daddr are host addresses that identify the specific testbed topology. These fields can be highly predictive within a single capture because the testbed assigns addresses, orders records, and schedules attacks reproducibly, but they carry no transport-level meaning that transfers to a new network. Retaining such fields produced within-dataset AUC near 1.000 but cross-dataset AUC near 0.300 in preliminary experiments. We applied the same exclusion principle to CICIoT2023 and TON_IoT: we removed flow IDs, IP addresses, port numbers, absolute timestamps, capture-file identifiers, and host identifiers before feature construction. We also excluded raw payload and packet-content features because they are not recoverable from all three flow exporters and would violate dimensional consistency.

3.4. Risk Bound and Prior Correction

Following Ben-David et al. [19], the target-domain risk admits the bound
e T h = e S h + d H P S , P T + λ ∗ + C n S , δ ,
in which dH denotes the divergence between source and target distributions, λ ∗ the error of the optimal joint hypothesis, and C a complexity term decreasing with source sample size nS. Section Testing the Divergence Term Empirically reports an empirical estimate of dH via maximum mean discrepancy for each transfer pair, along with its correlation with observed AUC, so the first component of this bound can be evaluated quantitatively rather than stated abstractly.
Prior shift contributes to the bound through its effect on the effective source risk. When a model trained under P S (Y) is evaluated under P T (Y), the mismatch is proportional to the Kullback–Leibler divergence between the two priors. For CICIoT2023 as source and TON_IoT as target, this divergence equals 0.61 nats, from which a substantial improvement from prior correction is predicted.

3.5. Per-Domain Normalization

We fitted a separate quantile transformer to each domain’s marginal distribution, mapping every feature to a standard Gaussian marginal. This procedure removes first-order covariate shift while requiring only unlabeled marginal observations from the target domain.

3.6. Per-Host Chronological Windowing

Flows were grouped by source host, ordered by timestamp, and partitioned into windows of length L = 8 with unit stride, never crossing host boundaries. On a controlled beacon domain, per-flow logistic regression attained AUC = 0.745, whereas a host-windowed GRU attained 0.999.

3.7. Domain-Invariant Training: An Investigated but Not Adopted Method

We evaluated V-REx [22] as a candidate domain-invariant training objective, using pseudo-environments formed by partitioning source hosts into groups of G hosts. Because only one source dataset is available per training run under the leave-one-pair-out protocol used here, we estimate per-environment risk variance within-dataset host groups rather than from genuinely independent environments, a departure from V-REx’s original multi-environment assumption that we flag explicitly rather than obscure. Across the full penalty-weight by environment-size sweep, empirical risk minimization achieved the highest cross-domain AUC (0.886) of any configuration tested; no V-REx setting improved on it in our data (Section 5.6). Consequently, we did not use V-REx to produce the pairwise results; we obtained them using standard empirical risk minimization throughout. We report the negative finding rather than omit it since it constrains which domain-generalization methods are worth applying under the small-number-of-source-domains regime typical of IoT intrusion detection. For future work with more than two source datasets, V-REx may be worth revisiting.

3.8. Calibration and Threshold Selection

A scalar temperature was fitted on source validation logits by minimizing the negative log-likelihood,
T ∗ = arg min T − 1 n ∑ i = 1 n y i log σ z i T + 1 − y i log 1 − σ z i T
where σ x = 1 1 + e − x is the sigmoid function, y i ∈ { 0 , 1 } are binary labels, z i are the logits (or inputs), and T > 0 is a temperature parameter.
We report two thresholds throughout: F1src, obtained at the threshold that maximizes F1 on source validation data and corresponds to zero-shot deployment where no target labels are available; and F1oracle, obtained by exhaustive grid search over the decision threshold using the target test labels. F1oracle is reported solely as a theoretical upper bound to isolate threshold miscalibration (Definition 5) from representation failure; it is not an operating point achievable without labeled target data, and no procedure in this paper claims otherwise. The gap between F1oracle and F1src measures the cost of using a source-calibrated threshold in the field.

3.9. Datasets

We selected three datasets to span three distinct flow exporters, three network environments, and a wide range of class priors (Table 2).
Table 2. Datasets employed in this study. Processed gives the exact number of rows read and used in all reported experiments.

3.9.1. CICIoT2023

Collection occurred across 105 heterogeneous IoT devices, containing 34 native labels (33 attack types plus BENIGN). Each native label maps to exactly one of seven canonical families, with no residual category; Table 3 summarizes the full mapping and per-label sample counts.
Table 3. CICIoT2023 native-label to canonical-family mapping, summarized by family.

3.9.2. TON_IoT

Collection occurred in a simulated industrial IoT environment, providing nine attack types with genuine Zeek-derived bidirectional flow statistics throughout.

3.9.3. Bot-IoT: Sampling Procedure and Exact Class Counts

Bot-IoT is distributed as four files of 1,000,000 rows each (4,000,000 rows total), consistent with the published 5% stratified sample used in most prior work citing this dataset. For computational tractability, we read 300,000 rows from each file (1,200,000 rows total), sequentially from the start of each file without further subsampling at this stage. Within this processed sample, 2017 rows are labeled benign (0.168%), and 1,197,983 are labeled attack (99.832%); the extreme imbalance reflects the dataset’s construction, in which the testbed was deliberately saturated with attack traffic. We applied a 60% prior correction by retaining all 2017 available benign flows and subsampling the attack class to 3025 flows without replacement, yielding a final training pool of 5042 flows. We report these absolute counts explicitly because the small number of benign flows is a material property of every Bot-IoT result in this paper: any transfer pair with Bot-IoT as source is trained on class-balanced information drawn from only 2017 unique benign observations.
We emphasize more prominently that every result in this paper with Bot-IoT as a source domain rests on only 2017 unique benign observations. The 60% prior correction retains all available benign flows and subsamples attacks to 3025, yielding a 5042-flow training pool. The benign class therefore includes 2017 unique flows, not 2017 independent benign hosts or sessions. This is a small effective sample for learning the benign manifold, and it is a material property of every Bot-IoT-as-source result reported here. Conclusions about Bot-IoT as a source domain should therefore be treated as provisional: the strong BOT→TON and BOT→CIC results may reflect the specific benign flows retained, and they may not generalize to other Bot-IoT subsamples or to operational Bot-IoT networks with different benign traffic volumes. We recommend that future work using Bot-IoT as a source report absolute benign counts and consider pooling benign traffic from additional captures.

3.10. Models and Evaluation Protocol

We compared three models under identical preprocessing: logistic regression with L2 regularization and balanced class weights; a random forest [34] of 100 trees, maximum depth 12, and balanced class weights, both implemented using scikit-learn [35]; and a two-layer GRU [36], a gated variant of the LSTM-style recurrent architecture [37], with hidden dimension 64, batch normalization, and sigmoid output, trained with Adam at learning rate 0.001 with early stopping. All three models, including logistic regression and random forest, were trained on an identical 20,000-flow cap drawn from the same stratified 85/15 train/validation split, so that no model in any reported comparison receives more training data than another; no result in this paper compares a capped model against an uncapped one. We evaluated every configuration with three independent random seeds.
We report performance as AUC, balanced accuracy, F1src, and F1oracle. We do not report accuracy: at the attack priors present in these datasets, a constant-positive predictor attains 97.8% accuracy on CICIoT2023, exceeding the accuracy of most published models evaluated on that dataset. We assessed differences between GRU and logistic regression using a paired t-test on the three seed-level AUC differences per pair, treating each independently retrained seed as one replicate; this is the statistically appropriate unit of replication for a claim about whether an effect would reproduce under retraining. McNemar’s test on pooled predictions and bootstrap confidence intervals were also computed and are reported but are not used to determine significance in the main text because pooling predictions across seeds before testing conflates test-set size with genuine cross-seed model agreement; Section Why Pooled and Seed-Level Significance Tests Can Disagree documents a case (CIC to TON) where the two approaches disagree and explains why the seed-level test is preferred.
All three models are trained on an identical 20,000-flow cap. This cap ensures fair comparison but is not claimed to be optimal. Larger training sets could change the relative ranking, particularly for TON→CIC, where logistic regression currently outperforms the GRU. With more data, the GRU’s higher capacity might allow it to overtake logistic regression; alternatively, if the limitation is representation or exporter-induced feature corruption rather than data volume, the simpler model may remain competitive. We therefore flag training set size as an open variable and recommend a size sweep for TON→CIC in future work.

3.11. Direct Test of the Flow-Exporter Mechanism

This subsection describes a controlled synthetic experiment designed to test the proposed mechanism directly: that CICFlowMeter’s SYN/ACK-based directional-feature heuristic destroys information in a way that measurably harms transfer, independent of any other property of CICIoT2023.
Synthetic flows were generated with fully known ground-truth forward and backward byte/packet counts, drawn from three traffic mechanisms with distinct directional signatures: near-symmetric benign TCP traffic (true forward fraction approximately 0.50), asymmetric TCP SYN-flood traffic (approximately 0.70 to 0.90), and highly asymmetric UDP/ICMP flood traffic (approximately 0.80 to 0.95). SYN and ACK packet counts were generated according to each mechanism’s real protocol behavior: for TCP mechanisms, SYN and ACK counts correlate with true forward/backward packet counts, as in a genuine handshake; for the UDP/ICMP mechanism, SYN and ACK counts are exactly zero, because these protocols carry no such flags. We then applied CICFlowMeter’s published heuristic to derive an approximated set of the eight directional canonical features and compared it against the ground truth.
We constructed two source domains with different mechanism mix compositions (simulating CICIoT2023-like and TON_IoT-like attack distributions) and one held-out target domain. Four quantities were then measured: (i) the approximation error and variance loss in the directional features, by mechanism; (ii) the downstream AUC of a classifier trained on the source domain’s true, counterfactual, Zeek-quality features and evaluated on the target domain’s true features; (iii) the downstream AUC of an otherwise identical classifier trained on the source domain’s CICFlowMeter-approximated features and evaluated on the same target; and (iv) the paired difference between (ii) and (iii), which isolates the causal cost of the heuristic alone, holding the classifier, training procedure, and cross-domain mechanism-mix shift fixed. We repeated this for both a linear classifier (logistic regression) and the GRU used elsewhere in this paper, across five and three random seeds, respectively.
The heuristic has no defined behavior for UDP or ICMP because these protocols carry no SYN or ACK flags. When both counts are zero, the ratio is undefined; CICFlowMeter’s implementation falls back to a clipped constant or a default value. Consequently, all UDP/ICMP flows in a window receive effectively the same directional-feature value regardless of their true forward/backward byte or packet counts. In our synthetic test, this collapsed within-mechanism variance to zero for the UDP/ICMP flood mechanism. This implies that datasets with substantial UDP/ICMP traffic—common in IoT (CoAP, MQTT over UDP, DNS, NTP, QUIC, and many DDoS/ICMP floods)—will have systematically degraded directional features for entire protocol classes. Models trained on such features cannot use directionality to separate benign from attack traffic within those classes, and any cross-dataset transfer involving such a source will inherit this limitation.

3.12. Confounding and Identification

Our design compares three datasets that differ simultaneously in flow exporter, traffic mix, device population, attack distribution, labeling procedure, and collection environment. Exporter identity is therefore perfectly confounded with dataset identity. The pairwise transfer matrix can show that exporter provenance is associated with transfer failure, but it cannot identify exporter as the sole or dominant cause. The synthetic experiment in Section 3.11 isolates one concrete mechanism, CICFlowMeter’s SYN/ACK-based directional-feature heuristic, and confirms that it is real and correctly signed, but its measured single-axis effect (+0.004 AUC under the GRU) is much smaller than the full empirical CIC→TON gap. We therefore interpret the evidence as follows: exporter provenance is a major, directly implicated contributor and a confounder that cross-dataset IoT IDS research must control for, but it is not established as the sole determinant. A fully isolated exporter test would require re-extracting features from the same packet captures with multiple exporters, which was not available for this study. We identify this as the highest-priority direction for future work.

4. Results

4.1. Pairwise Transfer Matrix

Table 4 reports GRU and logistic regression performance across all six directed transfer pairs; random forest is reported separately in Table 5 because its qualitative behavior differs sufficiently to warrant its own discussion.
Table 4. Pairwise transfer matrix, GRU vs. logistic regression (3 seeds, mean ± standard deviation). CIC = CICIoT2023; TON = TON_IoT; BOT = Bot-IoT. We assess significance with a paired t-test on the three seed-level AUC differences (df = 2). Note: The CIC→TON row uses the harmonized pipeline.
Table 5. Random forest, all six pairs (3 seeds, mean plus/minus standard deviation).
The CIC→TON row was originally reported under an intermediate preprocessing configuration (GRU AUC = 0.562 ± 0.096). That value is not directly comparable to the other rows. The values reported here are from the harmonized pipeline. Readers should therefore treat the CIC→TON GRU-vs-LR comparison as descriptive rather than inferential.
Four of six pairs reach significance in favor of the GRU under this seed-level test: CIC to BOT, TON to BOT, BOT to CIC, and BOT to TON. Two pairs do not: CIC to TON, where the sign of the difference is not even consistent across seeds (GRU favored in only 1 of 3), and TON to CIC, where all three seeds favor GRU but the magnitude varies too much (0.040 to 0.182) for three observations to establish significance. The GRU mean AUC across all pairs is 0.649; the logistic-regression mean is 0.529. Note that this corrected significant set differs from a naive combination of per-seed McNemar tests: CIC to BOT, sourced from CICFlowMeter, is significant, while TON to CIC, sourced from Zeek, is not. The exporter-quality pattern established in this paper therefore does not rest on the GRU-versus-logistic-regression significance pattern in Table 4; it rests primarily on the random forest balanced-accuracy collapse reported in Table 5.

Why Pooled and Seed-Level Significance Tests Can Disagree

An earlier version of this analysis reported significance by averaging per-seed McNemar p-values and per-seed bootstrap confidence interval bounds arithmetically. This procedure is not statistically valid and produced a misleading result for CIC to TON specifically: the three seeds individually gave McNemar p = 0.000, p = 0.000, and p = 1.000, with the AUC difference changing sign on the third seed (GRU trailed logistic regression by 0.070 and 0.045 AUC on the first two seeds, then led by 0.101 on the third). Averaging these three numbers produced p = 0.333 and a 95% interval of (−0.007, −0.002), an interval that excludes zero despite the underlying seeds not agreeing on which model is better. This apparent contradiction, an interval excluding zero alongside a non-significant p-value, is a direct symptom of averaging quantities that should never be averaged: p-values and confidence-interval bounds are not linear in the underlying evidence and averaging them across independent tests does not yield a valid combined test.
Instead, we computed two statistically coherent alternatives. The first pools every prediction from all three seeds into one large sample (for CIC to TON, n = 633,108) and runs a single McNemar test on the pooled predictions. This test has enormous power, driven by test-set size rather than genuine agreement between seeds, and returns chi-squared = 52,827, p < 0.000001, a result dominated by the two seeds where logistic regression clearly wins and largely insensitive to the third seed’s reversal. The second treats each seed as one independent replicate and runs a paired t-test on the three resulting AUC differences; with only three observations, this test has limited power by construction, but it correctly returns a non-significant result (p = 0.938) that reflects the genuine disagreement between seeds rather than masking it.
The seed-level paired test is adopted as the primary criterion throughout this paper because it answers the question a reader actually needs answered, namely whether retraining with a new random seed would likely reproduce the reported advantage, whereas the pooled test primarily reflects test-set size and can report high confidence in an effect that does not reliably replicate across independently trained models. We flag explicitly that three seeds give a paired test with only two degrees of freedom, which is a low-power regime; the qualitative conclusion for CIC to TON, namely that no reliable GRU advantage is established, would benefit from confirmation with additional seeds in future work.

4.2. Random Forest and the Exporter Effect Across Model Families

Table 5 reports random forest on the same six pairs. Random forest attains the highest mean AUC of the three classifiers (0.740), driven almost entirely by two pairs, TON to CIC (0.971) and TON to BOT (0.925), where it exceeds both the GRU and logistic regression by a wide margin. On the two CICFlowMeter-sourced pairs, however, its balanced accuracy collapses to 0.501 and 0.509, indistinguishable from a constant-positive predictor, despite an AUC on CIC to BOT (0.733) that in isolation would appear respectable. Unlike the GRU-versus-logistic-regression comparison in Table 4, this pattern directly compares magnitudes rather than testing significance, and it is therefore not subject to the seed-count and pooling issues; it is, correspondingly, the clearest and most robust evidence for the exporter effect reported in this paper.
Random forest partitions on absolute feature thresholds and has no sequential or recurrent structure whatsoever; the GRU processes ordered windows of flows. That both classifiers collapse specifically on CICFlowMeter-sourced pairs, and only on those pairs, rules out explanations grounded in recurrent-model inductive bias, sequence length, or architecture-specific overfitting. The shared factor across both model families is the source exporter.
Figure 1a visualizes the comparative performance of the three classifiers across all transfer pairs. This panel illustrates the central finding of this work: while random forest achieves the highest mean AUC, its performance depends heavily on the source dataset. The figure also shows that the collapse in balanced accuracy, our strongest evidence for the exporter effect, occurs exclusively on CICFlowMeter-sourced pairs. Figure 1b provides a detailed per-attack-family breakdown for the CICIoT2023-to-TON_IoT transfer, revealing a clear coverage threshold. Families with more than 1000 source samples (e.g., FLOOD) transfer with high F1, while those with fewer samples (e.g., WEB, BRUTEFORCE) fail completely, irrespective of the classifier used.
Figure 1. Cross-dataset transfer performance. (a) AUC for GRU, logistic regression, and random forest across all six directed pairs; error bars denote one standard deviation over three seeds; the source flow exporter is indicated beneath each pair. (b) Per-family F1 heatmap for the CICIoT2023 to TON_IoT pair; n/a denotes a family absent from the target dataset.

Testing the Divergence Term Empirically

Equation (1) predicts that transfer error grows with the H-divergence between source and target distributions, for which maximum mean discrepancy (MMD) is a standard computable proxy. We computed MMD for each pair in the quantile-normalized canonical feature space and correlated it with the observed GRU AUC (Table 6).
Table 6. Maximum mean discrepancy (RBF kernel) against observed GRU AUC, all six pairs.
With six transfer pairs, we lack the statistical power to draw strong conclusions about MMD’s predictive value. However, the observed correlation (r = −0.165) is inconsistent with the hypothesis that aggregate distributional distance alone explains the exporter effect, suggesting a structural rather than continuous mechanism. Aggregate distributional distance, whether computed over the whole feature space or restricted to exactly the features our mechanism hypothesis implicates, does not predict which pairs transfer successfully. The exporter effect is therefore not well described by a smooth divergence measure; it behaves as a discrete, structural property of which tool produced the source data rather than a continuous distance one could shrink incrementally. This motivates the direct mechanism test, which does not rely on an aggregate divergence proxy.

4.3. Comparison of Model Families on CICIoT2023 to TON_IoT

Table 7 reports all three model families on the single pair with complete multi-model results at a matched 20,000-flow training budget. Random forest attains the highest F1 (0.885) but a balanced accuracy of 0.581, again close to chance; the GRU attains the highest balanced accuracy (0.631).
Table 7. Model comparison on CICIoT2023 to TON_IoT, identical 20,000-flow training cap for all three models (3 seeds, mean plus/minus standard deviation).
The standard deviation of balanced accuracy for logistic regression (0.152) exceeds that of the other models by roughly a factor of two, reflecting sensitivity of the source-selected threshold to the validation partition on this especially difficult pair.

4.4. Per-Family Decomposition

Table 8 reports per-family F1 alongside source-domain training-sample counts (from Table 3). Table 8 does not support a universal 1000-sample coverage threshold. In the CIC→TON direction, BRUTEFORCE and WEB have very few post-correction source samples yet achieve high F1, while in the TON→CIC direction, several families with more than 1000 samples fail. Coverage is therefore necessary but not sufficient. Transfer success depends jointly on source coverage, target support, family separability, prior shift, and threshold calibration. We report the per-family counts and F1 values without imposing a threshold rule.
Table 8. Per-family F1 at the source-selected threshold, CICIoT2023 to TON_IoT, with source training set coverage from Table 3. The table is descriptive; it does not support a universal sample count threshold for transfer success.
Results revealed that RANSOMWARE is absent from CICIoT2023 yet attained F1 = 0.832 under transfer. Inspection of TON_IoT shows that most RANSOMWARE flows have zero recorded duration, consistent with synthetic traffic injection rather than a general property of ransomware traffic. This apparent zero-shot generalization is therefore flagged as a probable dataset-specific artifact rather than reported as evidence of genuine attack-family generalization; it should not be relied upon without corroboration from operational ransomware captures.

4.5. Ablation

Table 9 reports the ablation on CICIoT2023 to TON_IoT under the same three-seed, 20,000-flow-cap protocol used throughout this paper, so that ablation results are directly comparable to Table 4, Table 5 and Table 7 rather than resting on a separately configured single-seed run.
Table 9. Ablation, CICIoT2023 to TON_IoT (GRU, 3 seeds, mean plus/minus standard deviation, identical protocol to Table 4).
Under the consistent three-seed, capped-training protocol, per-seed standard deviations are substantial (up to plus/minus 0.18 AUC at the baseline configuration), and the final windowing step (E4 to E5, plus 0.008 AUC) is not distinguishable from zero given this variance. The two largest, statistically robust contributions are prior correction and scale-free feature engineering; the contribution of temporal windowing specifically, on this pair, cannot be established from the ablation alone at this sample size and should be read as inconclusive rather than as a positive finding. Independent support for the value of host-chronological windowing comes instead from the controlled beacon-domain result (AUC 0.745 to 0.999), where the effect is large enough to be visible without averaging over noisy real-data seeds.

Sensitivity to the Prior-Correction Target

We did not tune the 60% correction target used throughout this paper; Table 10 reports a sensitivity sweep over the plausible range.
Table 10. GRU AUC on CICIoT2023 to TON_IoT as a function of the training set prior-correction target (2 seeds, mean plus/minus standard deviation).
Performance is roughly flat across 50 to 70%, all substantially above the uncorrected 97.8–prior baseline (Table 9, row E1). No value in this range is distinguishable from the others given the observed variance; 60% was retained as a defensible interior choice, but the finding that matters is that some correction away from the native prior is necessary, not the specific target chosen.

4.6. Label Budget

On CIC→TON, the source-calibrated threshold is already well calibrated, so the zero-shot F1 (0.929) is higher than the recalibrated F1 at k = 50 (0.881). The label-budget procedure is designed for pairs where the source threshold is miscalibrated, i.e., where the gap between F1_oracle and F1_src is large. For CIC→TON, that gap is small, so the curve is flat or slightly lower. Fifty labeled target flows recover 98.7% of the k = 500 asymptotic value (0.881 vs. approximately 0.891) as shown in Table 11; they do not improve over the zero-shot source-calibrated threshold on this pair. Figure 2 should therefore be read as showing that a small label budget recovers most of the achievable recalibration gain when recalibration is needed, not that it universally improves zero-shot performance.
Table 11. Cross-dataset F1 as a function of labeled target-domain flows, k (CICIoT2023 to TON_IoT, calibrated GRU, means over 30 resamples per budget).
Figure 2. Cross-dataset F1 against labeled target-domain flows used for threshold selection. Markers denote means over 30 resamples; error bars denote one standard deviation. On CIC→TON, the source threshold is already near-optimal, so the curve is flat or slightly lower than zero-shot; the procedure is intended for pairs with a large F1_oracle − F1_src gap.
Figure 2 illustrates the substantial benefit of even a minimal amount of labeled target-domain data for threshold recalibration. The curve shows that the cross-dataset F1 performance on the CICIoT2023 to TON_IoT pair increases sharply with the first few labeled target flows and plateaus quickly. The figure confirms that a budget of approximately 50 labeled flows is sufficient to recover 98.7% of the asymptotic performance achievable with a much larger labeled set, making threshold recalibration a highly cost-effective strategy for improving zero-shot deployment performance.

4.7. Direct Test of the CICFlowMeter Mechanism

We ran the synthetic experiment described in Section 3.11 to test the exporter-quality hypothesis directly, rather than by association alone. Table 12 reports the directional-feature distortion; Table 13 reports the downstream classification cost. This experiment measures the isolated contribution of one mechanism. The full empirical gap of ≈0.3 AUC between CIC→TON and TON→BOT likely requires compounding across the eight approximated features and their interaction with the prior shift documented in Section 4.5. Testing the compounded effect requires re-extraction from real PCAPs, which was not available.
Table 12. Directional feature approximation error and variance collapse by synthetic traffic mechanism (source domain).
Table 13. Downstream classification cost of the heuristic, holding cross-domain mechanism mix shift, classifier, and training procedure fixed. Counterfactual trains on hypothetical Zeek-quality (true) source features; realistic trains on the actual CICFlowMeter-approximated source features. Both are evaluated on the same held-out target domain’s true features.
The heuristic reduces within-mechanism variance by one to four orders of magnitude, and for UDP/ICMP flood traffic, collapses to a value indistinguishable from zero: every synthetic UDP/ICMP flood flow receives an identical directional-feature value regardless of its true intensity because these protocols carry no SYN or ACK flags and the heuristic’s clip-bounded ratio therefore returns a constant. This confirms substantial mechanism-specific information destruction as a direct, measured quantity rather than an inference.
Under a linear classifier combined with per-domain quantile normalization, the heuristic’s cost is not distinguishable from zero: quantile normalization is rank-based, and the corrupted directional feature preserves the correct monotonic relationship to the attack label even though its magnitude is severely compressed, which is sufficient for a linear rank-sensitive classifier to recover most of the signal. Under the GRU, we observe a small but consistently positive cost (plus 0.0041 AUC, positive in all three seeds tested), confirming the mechanism is real, correctly signed, and detectable by a classifier of the same family used throughout this paper’s main results. The magnitude of this isolated, single-axis effect (approximately 0.004 AUC) is far smaller than the empirically observed CIC to TON gap from near-perfect within-dataset performance to near-random cross-dataset AUC. These results should therefore be read as proof of mechanism and as a lower bound on the exporter effect, not as a complete explanation of the empirical gap. The full CIC→TON degradation likely arises from compounding across multiple simultaneously corrupted canonical features, interaction with prior and concept shift, and dataset-level differences that the synthetic experiment necessarily simplifies away. We therefore report this as confirming evidence for one genuine contributing mechanism, not as a complete explanation of the empirical magnitude: the full effect most plausibly arises from the combined, compounding corruption of multiple canonical features simultaneously (this experiment isolated only the eight directional features), interaction with the concept- and prior-shift mechanisms documented in Section 3.1 and Section 4.5, and structure present in real CICIoT2023 traffic that a three-mechanism synthetic model necessarily simplifies away. Testing the multi-feature, compounded version of this experiment on re-extracted real packet captures is identified as the highest-priority direction for future work.

5. Discussion

5.1. Interpretation of the Exporter Effect

Correcting the significance-testing error described in Section Why Pooled and Seed-Level Significance Tests Can Disagree changes which specific pairs reach significance in the GRU-versus-logistic-regression comparison, but it does not change the balanced-accuracy evidence in Table 5, which was never subject to that error because it reports magnitudes rather than a pooled or averaged significance statistic. The exporter effect is best supported by three lines of evidence, each with a different robustness to confounding. As the most robust, random forest balanced accuracy collapses to 0.501 and 0.509 on exactly the two CICFlowMeter-sourced pairs and only those pairs, while remaining strong whenever Zeek or Argus supplied either dataset. Second, aggregate MMD does not predict which pairs transfer, indicating a discrete, tool-specific association rather than a smooth distributional one. Third, the controlled synthetic experiment confirms a small, correctly signed causal contribution from one concrete mechanism. Because exporter identity is confounded with dataset identity, we treat these as evidence that exporter provenance is a major contributor and a necessary control variable, not as proof that it is the exclusive cause.
Taken together, these results support a qualified rather than absolute version of the exporter-quality hypothesis. CICFlowMeter’s directional-feature approximation is a real, measurable, and causally implicated contributor to poor transfer, evidenced most clearly by the random forest balanced-accuracy collapse and by the direct mechanism test, but the GRU-versus-logistic-regression significance pattern should not be read as a fourth independent confirmation of a clean exporter-based rule; it is a noisier signal that partially agrees with the exporter pattern and partially does not, most plausibly because a two-classifier significance test with three seeds is underpowered relative to the balanced-accuracy magnitude comparison. The honest overall conclusion is that source flow-exporter identity is strongly associated with transfer failure through the balanced-accuracy channel (Section 4.1 and Section 4.2), that this association is not explained by aggregate distributional distance (Section Testing the Divergence Term Empirically), and that at least one concrete, testable component of the underlying mechanism has been directly confirmed at small effect size (Section 4.7), with the remaining gap most plausibly attributable to compounding across multiple simultaneously corrupted features and interaction with prior and concept shift.

5.2. Reporting Requirements at High Class Priors

Random forest illustrates a reporting problem with implications beyond this study. On CIC to TON, random forest attains F1 = 0.866 with balanced accuracy 0.501; on CIC to BOT, F1 = 0.749 with balanced accuracy 0.509. At these class priors, a constant-positive predictor attains comparable F1 values by construction. Reported in isolation, F1 value would suggest competent discrimination; balanced accuracy reveals the opposite. Because most IoT intrusion detection datasets exhibit attack priors above 0.7, F1 values reported without an accompanying prior-independent metric are difficult to interpret, and we recommend balanced accuracy as a mandatory companion metric for cross-dataset evaluation in this domain.

5.3. Coverage Requirements and Dataset Design

The per-family results (Table 8) do not support a universal 1000-sample coverage threshold. Some families with very few source samples transfer well, while some families with more than 1000 samples fail. Coverage is necessary but not sufficient; transfer depends jointly on coverage, target support, family separability, prior shift, and threshold calibration. Datasets intended to support cross-dataset evaluation should nonetheless provide adequate depth per family because sparse families are more likely to fail even when coverage alone does not determine the outcome.
CICIoT2023 provides 33 attack categories (Table 3), of which 3 (BRUTEFORCE, WEB, and BOTNET’s smallest constituent, BACKDOOR_MALWARE) fall below this threshold after prior correction. Datasets intended to support cross-dataset evaluation should provide sufficient depth per family or be distributed alongside companion captures sharing the same families at adequate sample counts.

5.4. Distinguishing Calibration Failure from Representation Failure

Reporting F1src and F1oracle separately (Section 3.8) distinguishes two failure modes conflated in prior literature. A high oracle F1 combined with a low deployed F1 indicates a correctly ranking score function and a misplaced operating point, remediable with a small, labeled sample (Section 4.6). Low values for both indicate representation failure. TON to BOT exhibits successful transfer with a small oracle gap; CIC to TON exhibits near-random AUC and is consistent with representation failure.

5.5. Recommendations

Based on the results discussed so far, the following recommendations are concluded:
  • Exporter selection: Zeek or Argus is preferable to CICFlowMeter for training data collection intended to support cross-environment deployment; this recommendation is now supported across two classifier families and a direct mechanism test, not by association alone.
  • Prior correction: Resample training sets to an attack rate of 50–65%; the sensitivity sweep (Table 10) shows that this is not critical to tune precisely, but it must be applied.
  • Labeling budget. Provision approximately 50 labeled target-domain flows for threshold recalibration when the source–target calibration gap is large. Where the source threshold is already near-optimal, the procedure recovers no additional performance and zero-shot deployment is preferable.
  • Metric reporting: Balanced accuracy should accompany coverage; it is necessary but not sufficient, and any F1 in all cross-dataset evaluations at class priors above 0.7.
  • Model selection is exporter-dependent, not universal: random forest is the strongest choice on Zeek- or Argus-sourced transfers (Table 5) but should not be used on CICFlowMeter-sourced training data, where its balanced accuracy is indistinguishable from chance despite a superficially reasonable AUC.
  • Statistical testing: At least three random seeds should be evaluated with paired significance testing; seed-level standard deviation observed here reached 0.18 AUC at some ablation configurations.

5.6. Domain-Invariant Training: A Negative Result

Section 3.7 describes V-REx as investigated but not adopted. Table 14 reports the full sweep for completeness.
Table 14. V-REx sweep, cross-domain AUC as a function of penalty weight beta and pseudo-environment size G (2 seeds times 3 target domains, mean). Dashes denote configurations not evaluated.
No configuration of the penalty weight or pseudo-environment size exceeded plain empirical risk minimization (beta = 0) on cross-domain AUC. Coarser pseudo-environments (larger G) substantially reduced the per-environment risk-variance estimate used internally by the penalty (from approximately 0.030 at G = 1 to approximately 0.004 at G = 20), narrowing the gap to ERM but not closing it. We interpret this as a genuine limitation of applying environment-variance penalties when only one or two source datasets are available, rather than a limitation specific to our implementation, and report it so that future work does not need to rediscover it.

5.7. Limitations

The seed-level significance tests in Table 4 use only three random seeds per pair, giving a paired t-test only two degrees of freedom. This is a low-power regime: the test correctly avoids the invalid combination procedure it replaces, but it can fail to detect real, consistent effects when between-seed variance is high relative to three observations, as for TON to CIC, where all three seeds agree on direction but the test does not reach significance. Five to ten seeds per pair would substantially strengthen every significance claim in this paper and is identified as a direct, low-effort extension for future work. The synthetic mechanism test isolates a single feature axis (the eight directional features) under a simplified three-mechanism traffic model; it demonstrates that the proposed mechanism is real and causally implicated but explicitly does not close the gap to the full empirical effect size, which likely requires testing joint corruption across multiple canonical features simultaneously, and ultimately, validation against re-extracted features from real packet captures rather than synthetic data, access to which was not available for this study. All three-model comparisons use a 20,000-flow training cap applied uniformly to logistic regression, random forest, and GRU; this was adopted so that no model receives a training data advantage over another, but larger training sets remain untested and could shift the relative ranking, particularly for the TON to CIC pair where logistic regression currently outperforms the GRU. The directional-feature approximation affects both CICIoT2023 and Bot-IoT; we excluded N-BaIoT because its packet-stream statistics cannot be mapped onto the flow-level canonical schema without loss. We tested V-REx only with host-grouped pseudo-environments drawn from a single source dataset per training run; we did not evaluate genuine multi-dataset environment pooling, which may behave differently.
A second limitation concerns the CIC→TON significance test. The original Table 4 reported this pair under an intermediate preprocessing configuration that is not comparable to the full pipeline used in Table 7 and Table 9. The CIC→TON GRU-vs-logistic-regression comparison should therefore be treated as descriptive only. This does not affect the paper’s central finding, which rests on the random forest balanced-accuracy collapse in Table 5—a magnitude comparison that does not depend on the affected significance test.
Exporter identity is perfectly confounded with dataset identity in this study. Differences in traffic mix, device population, attack distribution, labeling, and collection environment may also contribute to the observed transfer failures. The synthetic experiment isolates only one mechanism and measures a small single-axis effect; it does not close the full empirical gap. A fully isolated exporter test would require re-extracting features from the same packet captures with multiple exporters, which was not available here.

6. Conclusions

This study evaluated cross-dataset generalization for IoT intrusion detection across three real datasets generated by three distinct flow exporters, using a complete pairwise transfer matrix over three classifier families, logistic regression, random forest, and a GRU, each trained under an identical data budget and evaluated across three random seeds with paired significance testing.
No single classifier is uniformly best. Random forest attains the highest mean AUC (0.740), but its balanced accuracy collapses to near-chance specifically on the two transfer pairs sourced from CICFlowMeter, while remaining strong on Zeek- and Argus-sourced pairs. This exporter-dependent balanced-accuracy collapse is the paper’s central finding. A statistically valid, seed-level significance test finds the GRU significantly stronger than logistic regression on four of six pairs, but exporter identity alone does not cleanly determine which four; the balanced-accuracy result, not the classifier-significance pattern, is therefore the primary and most robust evidence for the exporter effect, and it is a stronger claim for resting on a magnitude comparison rather than a low-power three-seed significance test.
This association was tested directly rather than left as a plausible narrative. A controlled synthetic experiment confirms that CICFlowMeter’s SYN/ACK-based directional-feature heuristic destroys substantial, mechanism-specific information (variance reduced by up to four orders of magnitude for UDP/ICMP flood traffic) and produces a small but consistently and correctly signed transfer cost under a recurrent classifier, while an aggregate divergence measure (MMD) is shown not to predict transfer success at all, together indicating a structural, tool-specific effect rather than a smooth distributional one. The isolated single-mechanism effect measured here is smaller than the full empirical gap, and we report this honestly as a partial explanation rather than a complete one.
Correcting the training set class prior produced the largest robust ablation improvement under a consistent, adequately powered protocol; the specific correction target did not matter within a wide plausible range. Per-family analysis established a coverage threshold near 1000 source samples below which transfer fails regardless of classifier. On pairs where the source threshold is miscalibrated, fifty labeled target-domain flows recover most of the achievable recalibration gain; on pairs where the source threshold is already well calibrated, the procedure does not improve over zero-shot performance. We tested a domain-invariant training method (V-REx) and found it did not improve on standard training in this small-number-of-source-domains setting, a negative result reported to support subsequent work. Taken together, these findings indicate that flow exporter provenance is a major upstream confounder and a directly implicated contributor to cross-dataset performance in IoT intrusion detection. Because exporter identity is confounded with dataset identity in the present design, we do not claim that it is the sole determinant. Rather, we conclude that exporter provenance must be treated as a first-class control variable in cross-dataset IoT IDS research, alongside model choice and training set composition.

Author Contributions

Conceptualization, M.A.R. and M.A.; methodology, M.A.R.; software, M.A.R.; validation, M.A.R. and M.A.; formal analysis, M.A.R.; investigation, M.A.; data curation, M.A.; writing, original draft preparation, M.A.R.; writing, review and editing, M.A.; visualization, M.A.; supervision, M.A.R.; project administration, M.A.R. All authors have read and agreed to the published version of the manuscript.

Funding

The authors would like to thank the Research Chair of Prince Dr. Faisal bin Mishaal bin Saud bin Abdulaziz for Artificial Intelligence (CPFAI) at Qassim University for supporting this research work (QU-CPFAI-25EF192).

Data Availability Statement

The three datasets analyzed are publicly available. CICIoT2023 [https://www.unb.ca/cic/datasets/iotdataset-2023.html; accessed on 5 June 2026] is distributed by the Canadian Institute for Cybersecurity; TON_IoT [https://research.unsw.edu.au/projects/toniot-datasets; accessed on 8 June 2026] and Bot-IoT [https://research.unsw.edu.au/projects/bot-iot-dataset; accessed on 12 June 2026] are distributed by UNSW Canberra. The canonical feature adapters, evaluation pipeline, synthetic mechanism-test generator, and analysis scripts supporting the reported results are available from the corresponding author upon reasonable request.

Acknowledgments

During the preparation of this manuscript, the authors used Claude Pro (Opus 5) to assist with Python 3.11.15 code implementation for the evaluation pipeline and synthetic experiment, and Grammarly for language editing. The authors reviewed and tested all code. The authors take full responsibility for all content.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Al-Fuqaha, A.; Guizani, M.; Mohammadi, M.; Aledhari, M.; Ayyash, M. Internet of Things: A survey on enabling technologies, protocols, and applications. IEEE Commun. Surv. Tutor. 2015, 17, 2347–2376. [Google Scholar] [CrossRef] [Scilit]
  2. Yang, Y.; Wu, L.; Yin, G.; Li, L.; Zhao, H. A survey on security and privacy issues in Internet-of-Things. IEEE Internet Things J. 2017, 4, 1250–1258. [Google Scholar] [CrossRef] [Scilit]
  3. Diro, A.; Chilamkurti, N. Distributed attack detection scheme using deep learning approach for Internet of Things. Future Gener. Comput. Syst. 2018, 82, 761–768. [Google Scholar] [CrossRef] [Scilit]
  4. Restuccia, F.; D’Oro, S.; Melodia, T. Securing the Internet of Things in the age of machine learning and software-defined networking. IEEE Internet Things J. 2018, 5, 4829–4842. [Google Scholar] [CrossRef] [Scilit]
  5. Brun, O.; Yin, Y.; Gelenbe, E. Deep learning with dense random neural network for detecting attacks against IoT-connected home environments. Procedia Comput. Sci. 2018, 134, 458–463. [Google Scholar] [CrossRef] [Scilit]
  6. Sommer, R.; Paxson, V. Outside the closed world: On using machine learning for network intrusion detection. In Proceedings of the IEEE Symposium on Security and Privacy, Oakland, CA, USA, 16–19 May 2010; pp. 305–316. [Google Scholar]
  7. Ring, M.; Wunderlich, S.; Scheuring, D.; Landes, D.; Hotho, A. A survey of network-based intrusion detection data sets. Comput. Secur. 2019, 86, 147–167. [Google Scholar] [CrossRef] [Scilit]
  8. Apruzzese, G.; Colajanni, M.; Ferretti, L.; Guido, A.; Marchetti, M. On the effectiveness of machine and deep learning for cyber security. In Proceedings of the International Conference on Cyber Conflict, Tallinn, Estonia, 29 May–1 June 2018. [Google Scholar]
  9. Paxson, V. Bro: A system for detecting network intruders in real-time. Comput. Netw. 1999, 31, 2435–2463. [Google Scholar] [CrossRef] [Scilit]
  10. Lee, W.; Stolfo, S.J. A framework for constructing features and models for intrusion detection systems. ACM Trans. Inf. Syst. Secur. 2000, 3, 227–261. [Google Scholar] [CrossRef] [Scilit]
  11. Tang, T.A.; Mhamdi, L.; McLernon, D.; Zaidi, S.A.R.; Ghogho, M. Deep learning approach for network intrusion detection in software defined networking. In Proceedings of the International Conference on Wireless Networks and Mobile Communications, Fez, Morocco, 26–29 October 2016. [Google Scholar]
  12. Vinayakumar, R.; Alazab, M.; Soman, K.P.; Poornachandran, P.; Al-Nemrat, A.; Venkatraman, S. Deep learning approach for intelligent intrusion detection system. IEEE Access 2019, 7, 41525–41550. [Google Scholar] [CrossRef] [Scilit]
  13. Shone, N.; Ngoc, T.N.; Phai, V.D.; Shi, Q. A deep learning approach to network intrusion detection. IEEE Trans. Emerg. Top. Comput. Intell. 2018, 2, 41–50. [Google Scholar] [CrossRef] [Scilit]
  14. Mirsky, Y.; Doitshman, T.; Elovici, Y.; Shabtai, A. Kitsune: An ensemble of autoencoders for online network intrusion detection. In Proceedings of the Network and Distributed System Security Symposium, San Diego, CA, USA, 18–21 February 2018. [Google Scholar]
  15. Meidan, Y.; Bohadana, M.; Mathov, Y.; Mirsky, Y.; Shabtai, A.; Breitenbacher, D.; Elovici, Y. N-BaIoT: Network-based detection of IoT botnet attacks using deep autoencoders. IEEE Pervasive Comput. 2018, 17, 12–22. [Google Scholar] [CrossRef] [Scilit]
  16. Diro, A.; Chilamkurti, N. Leveraging LSTM networks for attack detection in fog-to-things communications. IEEE Commun. Mag. 2018, 56, 124–130. [Google Scholar] [CrossRef] [Scilit]
  17. Popoola, S.I.; Adebisi, B.; Ande, R.; Hammoudeh, M.; Anoh, K.; Atayero, A.A. SMOTE-DRNN: A deep learning algorithm for botnet detection in the Internet-of-Things networks. Sensors 2021, 21, 2985. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Sarhan, M.; Layeghy, S.; Moustafa, N.; Portmann, M. NetFlow datasets for machine learning-based network intrusion detection systems. In Big Data Technologies and Applications; Springer: Cham, Switzerland, 2021; pp. 117–135. [Google Scholar]
  19. Ben-David, S.; Blitzer, J.; Crammer, K.; Kulesza, A.; Pereira, F.; Vaughan, J.W. A theory of learning from different domains. Mach. Learn. 2010, 79, 151–175. [Google Scholar] [CrossRef] [Scilit]
  20. Sun, B.; Saenko, K. Deep CORAL: Correlation alignment for deep domain adaptation. In Proceedings of the European Conference on Computer Vision Workshops, Amsterdam, The Netherlands, 8–16 October 2016; pp. 443–450. [Google Scholar]
  21. Ganin, Y.; Ustinova, E.; Ajakan, H.; Germain, P.; Larochelle, H.; Laviolette, F.; Marchand, M.; Lempitsky, V. Domain-adversarial training of neural networks. J. Mach. Learn. Res. 2016, 17, 2030–2096. [Google Scholar]
  22. Krueger, D.; Caballero, E.; Jacobsen, J.-H.; Zhang, A.; Binas, J.; Zhang, D.; Le Priol, R.; Courville, A. Out-of-distribution generalization via risk extrapolation. In Proceedings of the International Conference on Machine Learning, Virtual, 18–24 July 2021. [Google Scholar]
  23. Sagawa, S.; Koh, P.W.; Hashimoto, T.B.; Liang, P. Distributionally robust neural networks for group shifts. In Proceedings of the International Conference on Learning Representations, Addis Ababa, Ethiopia, 26–30 April 2020. [Google Scholar]
  24. Draper-Gil, G.; Lashkari, A.H.; Mamun, M.S.I.; Ghorbani, A.A. Characterization of encrypted and VPN traffic using time-related features. In Proceedings of the International Conference on Information Systems Security and Privacy, Rome, Italy, 19–21 February 2016. [Google Scholar]
  25. Zeek Project. The Zeek Network Security Monitor. Available online: https://zeek.org (accessed on 3 September 2026).
  26. QoSient LLC. Argus: Audit Record Generation and Utilization System. Available online: https://openargus.org (accessed on 3 September 2026).
  27. Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K.Q. On calibration of modern neural networks. In Proceedings of the International Conference on Machine Learning, Sydney, Australia, 6–11 August 2017; pp. 1321–1330. [Google Scholar]
  28. Saerens, M.; Latinne, P.; Decaestecker, C. Adjusting the outputs of a classifier to new a priori probabilities: A simple procedure. Neural Comput. 2002, 14, 21–41. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Lipton, Z.C.; Wang, Y.-X.; Smola, A. Detecting and correcting for label shift with black box predictors. In Proceedings of the International Conference on Machine Learning, Stockholm, Sweden, 10–15 July 2018. [Google Scholar]
  30. Gretton, A.; Borgwardt, K.M.; Rasch, M.J.; Scholkopf, B.; Smola, A. A kernel two-sample test. J. Mach. Learn. Res. 2012, 13, 723–773. [Google Scholar]
  31. Neto, E.C.P.; Dadkhah, S.; Ferreira, R.; Zohourian, A.; Lu, R.; Ghorbani, A.A. CICIoT2023: A real-time dataset and benchmark for large-scale attacks in IoT environment. Sensors 2023, 23, 5941. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  32. Alsaedi, A.; Moustafa, N.; Tari, Z.; Mahmood, A.; Anwar, A. TON_IoT telemetry dataset: A new generation dataset of IoT and IIoT for data-driven intrusion detection systems. IEEE Access 2020, 8, 165130–165150. [Google Scholar] [CrossRef] [Scilit]
  33. Koroniotis, N.; Moustafa, N.; Sitnikova, E.; Turnbull, B. Towards the development of realistic botnet dataset in the Internet of Things for network forensic analytics: Bot-IoT dataset. Future Gener. Comput. Syst. 2019, 100, 779–796. [Google Scholar] [CrossRef] [Scilit]
  34. Breiman, L. Random forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef] [Scilit]
  35. Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; et al. Scikit-learn: Machine learning in Python. J. Mach. Learn. Res. 2011, 12, 2825–2830. [Google Scholar]
  36. Cho, K.; van Merrienboer, B.; Gulcehre, C.; Bahdanau, D.; Bougares, F.; Schwenk, H.; Bengio, Y. Learning phrase representations using RNN encoder-decoder for statistical machine translation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Doha, Qatar, 25–29 October 2014; pp. 1724–1734. [Google Scholar]
  37. Hochreiter, S.; Schmidhuber, J. Long short-term memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Article metric data becomes available approximately 24 hours after publication online.