Figure 1.
System architecture showing the network model, passive attacker, and AP-level cover traffic injection. Real frames (solid lines) and dummy frames (dashed lines) are indistinguishable to the attacker. Dummy frames are discarded by the AP before forwarding to the internet.
Figure 1.
System architecture showing the network model, passive attacker, and AP-level cover traffic injection. Real frames (solid lines) and dummy frames (dashed lines) are indistinguishable to the attacker. Dummy frames are discarded by the AP before forwarding to the internet.
Figure 2.
Overview of the methodology. A trace-level, AP-assisted cover-traffic injection mechanism is evaluated against a passive metadata-fingerprinting attacker and converted into an empirical zero-sum game. The workflow spans the system and threat model, defender obfuscation, the windowed feature model, attacker evaluation, and game-theoretic analysis, scaling four devices to 198 scenario rows (156 unique defender configurations, since donor injection is -independent) and eight attacker classifiers, yielding a 198 × 8 empirical payoff matrix of pairing-unaware balanced accuracy.
Figure 2.
Overview of the methodology. A trace-level, AP-assisted cover-traffic injection mechanism is evaluated against a passive metadata-fingerprinting attacker and converted into an empirical zero-sum game. The workflow spans the system and threat model, defender obfuscation, the windowed feature model, attacker evaluation, and game-theoretic analysis, scaling four devices to 198 scenario rows (156 unique defender configurations, since donor injection is -independent) and eight attacker classifiers, yielding a 198 × 8 empirical payoff matrix of pairing-unaware balanced accuracy.
Figure 3.
Data processing pipeline from raw 802.11 captures through cover traffic injection, feature extraction, cross-validated classification, and game-theoretic analysis. The pipeline produces 1584 classifier–configuration pairs (198 defender configurations × 8 attacker classifiers), each evaluated with 10-fold GroupKFold cross-validation, yielding 15,840 total model fits.
Figure 3.
Data processing pipeline from raw 802.11 captures through cover traffic injection, feature extraction, cross-validated classification, and game-theoretic analysis. The pipeline produces 1584 classifier–configuration pairs (198 defender configurations × 8 attacker classifiers), each evaluated with 10-fold GroupKFold cross-validation, yielding 15,840 total model fits.
Figure 4.
Baseline confusion matrices for all eight classifiers at s. Values represent normalized per-class recall. Misclassification is predominantly pairwise within donor pairs, Bulb↔Plug, while Camera and Doorbell are classified with near-perfect accuracy.
Figure 4.
Baseline confusion matrices for all eight classifiers at s. Values represent normalized per-class recall. Misclassification is predominantly pairwise within donor pairs, Bulb↔Plug, while Camera and Doorbell are classified with near-perfect accuracy.
Figure 5.
Comparison of accuracy, balanced accuracy, and macro-F1 across four scenarios: baseline ( s), donor ( s, ), donor ( s, ), and donor ( s, ). All three metrics are consistent under baseline conditions and collapse together under donor injection, confirming that the defense effectiveness is not an artifact of metric choice.
Figure 5.
Comparison of accuracy, balanced accuracy, and macro-F1 across four scenarios: baseline ( s), donor ( s, ), donor ( s, ), and donor ( s, ). All three metrics are consistent under baseline conditions and collapse together under donor injection, confirming that the defense effectiveness is not an artifact of metric choice.
Figure 6.
Attacker balanced accuracy by obfuscation method at the equilibrium-relevant window s: (a) Fixed, (b) Exponential, (c) Uniform, (d) Donor. Rows are attacker classifiers; columns are injection intensities . Synthetic methods (a–c) leave the ensemble attackers (RF, XGB, ET) at or above 0.87 at every and no classifier below 0.60, while donor (d) drives the ensembles to 0.00 and holds every classifier at or below 0.26. The identical columns in (d) reflect the donor method’s -independence; the same qualitative pattern holds at all six window sizes.
Figure 6.
Attacker balanced accuracy by obfuscation method at the equilibrium-relevant window s: (a) Fixed, (b) Exponential, (c) Uniform, (d) Donor. Rows are attacker classifiers; columns are injection intensities . Synthetic methods (a–c) leave the ensemble attackers (RF, XGB, ET) at or above 0.87 at every and no classifier below 0.60, while donor (d) drives the ensembles to 0.00 and holds every classifier at or below 0.26. The identical columns in (d) reflect the donor method’s -independence; the same qualitative pattern holds at all six window sizes.
Figure 7.
Donor confusion matrices for all eight classifiers at s and . Values represent normalized per-class recall.
Figure 7.
Donor confusion matrices for all eight classifiers at s and . Values represent normalized per-class recall.
Figure 8.
Interaction between window size and obfuscation method for each classifier. Lines represent obfuscation methods, and shaded regions indicate standard deviation across cross-validation folds. Synthetic methods generally show increasing balanced accuracy with window size, whereas the donor method generally shows decreasing balanced accuracy.
Figure 8.
Interaction between window size and obfuscation method for each classifier. Lines represent obfuscation methods, and shaded regions indicate standard deviation across cross-validation folds. Synthetic methods generally show increasing balanced accuracy with window size, whereas the donor method generally shows decreasing balanced accuracy.
Figure 9.
ROC curves comparing baseline ( s) and donor ( s, ) conditions. AUC drops from 0.946–0.990 under baseline to 0.548–0.758 under donor injection, with KNN falling to near-random (0.548) and MLP retaining the highest residual AUC (0.758). The dashed diagonal line indicates the performance of a random classifier.
Figure 9.
ROC curves comparing baseline ( s) and donor ( s, ) conditions. AUC drops from 0.946–0.990 under baseline to 0.548–0.758 under donor injection, with KNN falling to near-random (0.548) and MLP retaining the highest residual AUC (0.758). The dashed diagonal line indicates the performance of a random classifier.
Figure 10.
Per-device recall distributions across all obfuscation regimes. Bimodal distributions reflect the separation between synthetic configurations, high recall, right cluster, and donor configurations, low recall, left cluster. Baseline recall is indicated by reference lines.
Figure 10.
Per-device recall distributions across all obfuscation regimes. Bimodal distributions reflect the separation between synthetic configurations, high recall, right cluster, and donor configurations, low recall, left cluster. Baseline recall is indicated by reference lines.
Figure 11.
Pareto frontier of privacy versus defender cost. Privacy is defined as , where is the attacker’s balanced accuracy against defender configuration i using classifier j. Defender cost is measured as the mean injection gap (log scale); smaller gaps imply denser injection. Synthetic methods cluster in the low-privacy region across the cost range, while donor methods occupy a distinct high-privacy region. No synthetic configuration is Pareto-optimal.
Figure 11.
Pareto frontier of privacy versus defender cost. Privacy is defined as , where is the attacker’s balanced accuracy against defender configuration i using classifier j. Defender cost is measured as the mean injection gap (log scale); smaller gaps imply denser injection. Synthetic methods cluster in the low-privacy region across the cost range, while donor methods occupy a distinct high-privacy region. No synthetic configuration is Pareto-optimal.
Figure 12.
Aggregated attacker balanced accuracy by the obfuscation method and window size, averaged across all classifiers and values. The donor method, bottom row, achieves dramatically lower accuracy than all synthetic methods and baseline across every window size, confirming that behavioral mimicry is the decisive factor regardless of observation duration.
Figure 12.
Aggregated attacker balanced accuracy by the obfuscation method and window size, averaged across all classifiers and values. The donor method, bottom row, achieves dramatically lower accuracy than all synthetic methods and baseline across every window size, confirming that behavioral mimicry is the decisive factor regardless of observation duration.
Table 1.
Comparison of related traffic obfuscation defenses against IoT device fingerprinting.
Table 1.
Comparison of related traffic obfuscation defenses against IoT device fingerprinting.
| Work | Year | Defense Type | Layer | Features | Classifiers | Game Theory | Volume Analysis |
|---|
| [1] | 2022 | Paired-device shaping | MAC | 2 | 3 | No | No |
| [16] | 2023 | Random segmentation | MAC | Size-based | 4 | No | No |
| [10] | 2022 | Eight padding methods | MAC | Full size dist. | 1 (RF) | No | No |
| [17] | 2019 | Stochastic padding | Network | Traffic rates | Theoretical | No | Partial |
| [18] | 2022 | Dynamic dummy traffic | Network | Timing | 2 | No | No |
| [19] | 2022 | DP-based shaping | Network | Mixed | 3 | No | No |
| [11] | 2025 | GAN-based injection | Network | Mixed | 3 | No | No |
| [24] | 2024 | Moving target defense | System | RF fingerprint | RL-based | Partial (RL) | No |
| This work | 2026 | Donor mimicry + 3 synthetic | MAC | 18 features | 8 (all families) | Full NE (198 × 8) | Yes |
Table 2.
Device characteristics.
Table 2.
Device characteristics.
| Device | MAC Address | Packets | Duration | Median IAT | Idle % | Donor |
|---|
| Bulb | 48:e1:e9:1a:22:6c | 459 | 1784 s | 0.115 s | 89.1% | Plug |
| Plug | d8:47:32:c2:24:be | 50 | 1605 s | 0.260 s | 98.7% | Bulb |
| Doorbell | 9c:8e:cd:27:1f:3b | 218,551 | 1807 s | 0.003 s | 0.1% | Camera |
| Camera | 2c:aa:8e:8f:74:2d | 630,907 | 1820 s | 0.0004 s | 17.5% | Doorbell |
Table 3.
Feature definitions and categories for the 18 statistical features extracted per observation window.
Table 3.
Feature definitions and categories for the 18 statistical features extracted per observation window.
| Feature | Mathematical Definition | Unit | Category |
|---|
| total_packets | , count of packets in window | count | size |
| total_bytes | , sum of packet lengths | bytes | size |
| unique_sizes | , count of distinct packet sizes | count | size |
| avg_pkt_size | , mean packet length | bytes | size |
| std_pkt_size | | bytes | size |
| max_pkt_size | , maximum packet length | bytes | size |
| mode_pkt_size | , most frequent size | bytes | size |
| var_pkt_size | , variance of sizes | bytes2 | size |
| var_size_freq | Variance of size-frequency counts | count2 | size |
| mean_iat | Mean inter-arrival time | seconds | timing |
| median_iat | Median IAT | seconds | timing |
| var_iat | Variance of IATs | s2 | timing |
| std_iat | Standard deviation of IATs | seconds | timing |
| skew_iat | Skewness of IATs | — | timing |
| kurt_iat | Excess kurtosis of IATs | — | timing |
| min_iat | Minimum IAT | seconds | timing |
| max_iat | Maximum IAT | seconds | timing |
| bytes_per_sec | total_bytes / window duration | bytes/s | rate |
Table 4.
Zero-window prevalence under baseline (no obfuscation) conditions. Values shown as zero-windows/total-windows (percentage).
Table 4.
Zero-window prevalence under baseline (no obfuscation) conditions. Values shown as zero-windows/total-windows (percentage).
| Device | s | s | s | s |
|---|
| Bulb | 1591/1785 (89.1%) | 57/180 (31.7%) | 1/31 (3.2%) | 1/16 (6.2%) |
| Plug | 1585/1606 (98.7%) | 144/162 (88.9%) | 14/28 (50.0%) | 2/15 (13.3%) |
| Doorbell | 3/1808 (0.2%) | 1/182 (0.5%) | 1/32 (3.1%) | 1/17 (5.9%) |
| Camera | 319/1821 (17.5%) | 1/183 (0.5%) | 1/32 (3.1%) | 1/17 (5.9%) |
Table 5.
Hyperparameter specifications and attacker rationale per classifier.
Table 5.
Hyperparameter specifications and attacker rationale per classifier.
| Model | Family | Hyperparameters | Attacker Rationale |
|---|
| Random Forest | Classical Ensemble | n_estimators=200, max_depth=None, min_samples_leaf=2 | Widely used baseline ensemble; robust to overfitting. |
| XGBoost | Gradient Boosting | n_estimators=200, max_depth=6, lr=0.1, subsample=0.8 | State-of-the-art gradient boosting; top performer on tabular data. |
| KNN | Instance-Based | k=5, weights=distance, metric=minkowski | Non-parametric; captures local decision boundaries. |
| SVM-RBF | Kernel Method | kernel=rbf, C=10.0, gamma=scale | Powerful nonlinear classifier via RBF kernel. |
| Logistic Regression | Linear Model | C=1.0, solver=lbfgs, max_iter=1000 | Linear baseline; tests linear separability. |
| MLP | Neural Network | layers=(128,64), relu, early_stopping | Feedforward NN; tests deep nonlinear representations. |
| Gaussian NB | Probabilistic | (no hyperparameters) | Naive Bayes; extremely fast baseline. |
| Extra-Trees | Randomized Ensemble | n_estimators=200, max_depth=None, min_samples_leaf=2 | More randomized than RF; often lower variance. |
Table 6.
Class balance: number of observation windows per device at representative window sizes under baseline conditions.
Table 6.
Class balance: number of observation windows per device at representative window sizes under baseline conditions.
| Window | Bulb | Camera | Doorbell | Plug |
|---|
| 1 s | 25.4% (1785) | 25.9% (1821) | 25.8% (1808) | 22.9% (1606) |
| 10 s | 25.5% (180) | 25.9% (183) | 25.7% (182) | 22.9% (162) |
| 60 s | 25.2% (31) | 26.0% (32) | 26.0% (32) | 22.8% (28) |
| 120 s | 24.6% (16) | 26.2% (17) | 26.2% (17) | 23.1% (15) |
Table 7.
Baseline balanced accuracy (mean ± std over 10-fold CV) for all classifiers across observation window sizes. Bold indicates the highest accuracy per column.
Table 7.
Baseline balanced accuracy (mean ± std over 10-fold CV) for all classifiers across observation window sizes. Bold indicates the highest accuracy per column.
| Classifier | s | s | s | s | s | s |
|---|
| Random Forest | | | | | | |
| XGBoost | | | | | | |
| KNN | | | | | | |
| SVM-RBF | | | | | | |
| Logistic Reg. | | | | | | |
| MLP | | | | | | |
| Gaussian NB | | | | | | |
| Extra-Trees | | | | | | |
Table 8.
Controlled volume comparison ( s).
Table 8.
Controlled volume comparison ( s).
| Configuration | Injected Packets | Ratio | Best Attacker Acc. |
|---|
| Donor (any ) | 852,855 | 1.0× | 0.335 |
| Fixed (matched) | 707,000 | 0.8× | 0.939 |
| Exponential (matched) | 743,953 | 0.9× | 0.939 |
| Uniform (matched) | 706,614 | 0.8× | 0.939 |
| Fixed (8× more) | 7,070,510 | 8.3× | 0.952 |
| Fixed (83× more) | 70,772,621 | 83× | 0.952 |
Table 9.
Feature importance at observation window s for the three equilibrium-relevant classifiers (RF, GNB, MLP), under baseline (BL) and donor (DN) conditions. RF uses impurity-based Gini importance; GNB and MLP use permutation importance on test folds (10 repeats per fold, averaged across 10 GroupKFold folds, scored by balanced accuracy). Each column is normalized to sum to 1.0 across the 18 features.
Table 9.
Feature importance at observation window s for the three equilibrium-relevant classifiers (RF, GNB, MLP), under baseline (BL) and donor (DN) conditions. RF uses impurity-based Gini importance; GNB and MLP use permutation importance on test folds (10 repeats per fold, averaged across 10 GroupKFold folds, scored by balanced accuracy). Each column is normalized to sum to 1.0 across the 18 features.
| Feature | Category | BL_RF | DN_RF | BL_GNB | DN_GNB | BL_MLP | DN_MLP |
|---|
| total_packets | Size | 0.0949 | 0.0582 | 0.0697 | 0.0316 | 0.0029 | 0.0287 |
| total_bytes | Size | 0.0814 | 0.0727 | 0.0742 | 0.0409 | 0 | 0.0211 |
| unique_sizes | Size | 0.0726 | 0.0590 | 0.0616 | 0.0323 | 0.0029 | 0.0236 |
| avg_pkt_size | Size | 0.0137 | 0.0810 | 0.0422 | 0.0700 | 0.0039 | 0.0496 |
| std_pkt_size | Size | 0.1052 | 0.0786 | 0.0616 | 0.0367 | 0.0828 | 0.0558 |
| max_pkt_size | Size | 0.0729 | 0.0244 | 0.0606 | 0.0355 | 0.0955 | 0.0306 |
| mode_pkt_size | Size | 0.0351 | 0.0182 | 0.0564 | 0.0438 | 0.0292 | 0.1541 |
| var_pkt_size | Size | 0.1078 | 0.0872 | 0.0616 | 0.0552 | 0.0799 | 0.0556 |
| var_size_freq | Size | 0.0692 | 0.0797 | 0.1054 | 0.0604 | 0 | 0.0333 |
| mean_iat | Timing | 0.0418 | 0.0679 | 0.0609 | 0.0917 | 0.1846 | 0.1516 |
| median_iat | Timing | 0.0380 | 0.0266 | 0.0123 | 0.0389 | 0.0341 | 0.0554 |
| var_iat | Timing | 0.0269 | 0.0445 | 0.0761 | 0.0624 | 0.0268 | 0.0636 |
| std_iat | Timing | 0.0294 | 0.0550 | 0.0604 | 0.0955 | 0.0166 | 0.1274 |
| skew_iat | Timing | 0.0511 | 0.0290 | 0.0126 | 0.0507 | 0.0429 | 0.0040 |
| kurt_iat | Timing | 0.0312 | 0.0295 | 0.0586 | 0.0571 | 0.0205 | 0.0020 |
| min_iat | Timing | 0.0114 | 0.0624 | 0.0214 | 0.0643 | 0.0044 | 0.0050 |
| max_iat | Timing | 0.0259 | 0.0414 | 0.0302 | 0.0922 | 0.3731 | 0.0938 |
| bytes_per_sec | Rate | 0.0917 | 0.0849 | 0.0742 | 0.0409 | 0 | 0.0449 |
| Size (sum) | — | 0.6528 | 0.5590 | 0.5932 | 0.4064 | 0.2971 | 0.4524 |
| Timing (sum) | — | 0.2555 | 0.3562 | 0.3326 | 0.5527 | 0.7029 | 0.5027 |
| Rate (sum) | — | 0.0917 | 0.0849 | 0.0742 | 0.0409 | 0 | 0.0449 |
Table 10.
Feature importance at observation window
s, under baseline (BL) and donor (DN) conditions. Columns and normalization as in
Table 9. Within-window method effects are reported separately in
Table 11.
Table 10.
Feature importance at observation window
s, under baseline (BL) and donor (DN) conditions. Columns and normalization as in
Table 9. Within-window method effects are reported separately in
Table 11.
| Feature | Category | BL_RF | DN_RF | BL_GNB | DN_GNB | BL_MLP | DN_MLP |
|---|
| total_packets | Size | 0.0700 | 0.0564 | 0.0740 | 0.0123 | 0 | 0.0828 |
| total_bytes | Size | 0.0708 | 0.0742 | 0.0812 | 0.0102 | 0 | 0.0438 |
| unique_sizes | Size | 0.0797 | 0.0389 | 0.0551 | 0.0596 | 0.0138 | 0.0540 |
| avg_pkt_size | Size | 0.0144 | 0.0778 | 0.0399 | 0.1190 | 0 | 0.0590 |
| std_pkt_size | Size | 0.0939 | 0.0737 | 0.0551 | 0.1190 | 0.1451 | 0.0480 |
| max_pkt_size | Size | 0.0782 | 0.0266 | 0.0551 | 0.0964 | 0.1919 | 0.0301 |
| mode_pkt_size | Size | 0.0136 | 0.0090 | 0.0524 | 0.0327 | 0.0133 | 0.0581 |
| var_pkt_size | Size | 0.1024 | 0.0801 | 0.0551 | 0.0273 | 0.1478 | 0.0583 |
| var_size_freq | Size | 0.0448 | 0.0769 | 0.0971 | 0.0155 | 0 | 0.0466 |
| mean_iat | Timing | 0.0564 | 0.0613 | 0.0524 | 0.0857 | 0 | 0.0673 |
| median_iat | Timing | 0.0312 | 0.0417 | 0.0091 | 0.0482 | 0.0083 | 0.0560 |
| var_iat | Timing | 0.0564 | 0.0574 | 0.0927 | 0.0675 | 0.0557 | 0.0696 |
| std_iat | Timing | 0.0557 | 0.0694 | 0.0576 | 0.0985 | 0.0041 | 0.0611 |
| skew_iat | Timing | 0.0707 | 0.0291 | 0.0090 | 0.0289 | 0.0908 | 0.0811 |
| kurt_iat | Timing | 0.0330 | 0.0329 | 0.0697 | 0.0375 | 0.0248 | 0.0109 |
| min_iat | Timing | 0.0055 | 0.0463 | 0.0220 | 0.0653 | 0.0309 | 0.0705 |
| max_iat | Timing | 0.0588 | 0.0639 | 0.0413 | 0.0664 | 0.2599 | 0.0393 |
| bytes_per_sec | Rate | 0.0644 | 0.0847 | 0.0812 | 0.0102 | 0.0138 | 0.0635 |
| Size (sum) | — | 0.5679 | 0.5135 | 0.5650 | 0.4919 | 0.5118 | 0.4806 |
| Timing (sum) | — | 0.3676 | 0.4018 | 0.3538 | 0.4979 | 0.4744 | 0.4559 |
| Rate (sum) | — | 0.0644 | 0.0847 | 0.0812 | 0.0102 | 0.0138 | 0.0635 |
Table 11.
Within-window method effects on category-level feature importance. Each row reports the donor-minus-baseline change in the Size, Timing, and Rate importance aggregates at a fixed window size, computed from the columns of
Table 9 and
Table 10. Positive values indicate importance gained under donor injection; negative values indicate importance lost. RF and GNB lose size importance and gain timing importance at both windows; MLP shows a timing-to-size shift at
s and a near-null shift at
s.
Table 11.
Within-window method effects on category-level feature importance. Each row reports the donor-minus-baseline change in the Size, Timing, and Rate importance aggregates at a fixed window size, computed from the columns of
Table 9 and
Table 10. Positive values indicate importance gained under donor injection; negative values indicate importance lost. RF and GNB lose size importance and gain timing importance at both windows; MLP shows a timing-to-size shift at
s and a near-null shift at
s.
| Classifier | Window | Size | Timing | Rate |
|---|
| RF | s | | | |
| RF | s | | | |
| GNB | s | | | |
| GNB | s | | | |
| MLP | s | | | |
| MLP | s | | | |
Table 12.
Marginal decomposition of attacker balanced accuracy and macro-F1 by experimental factor. Each row averages across all other factors. The dispersion column is the cell-to-cell standard deviation across the marginal, not the per-cell cross-validation standard deviation.
Table 12.
Marginal decomposition of attacker balanced accuracy and macro-F1 by experimental factor. Each row averages across all other factors. The dispersion column is the cell-to-cell standard deviation across the marginal, not the per-cell cross-validation standard deviation.
| Factor | Configuration | Bal. Accuracy | Macro-F1 | Number of Cells |
|---|
| Method | Baseline | | | 48 |
| | Fixed | | | 384 |
| | Exponential | | | 384 |
| | Uniform | | | 384 |
| | Donor | | | 384 |
| Window size | 1 s | | | 264 |
| | 5 s | | | 264 |
| | 10 s | | | 264 |
| | 30 s | | | 264 |
| | 60 s | | | 264 |
| | 120 s | | | 264 |
| Classifier | Random Forest | | | 198 |
| | XGBoost | | | 198 |
| | KNN | | | 198 |
| | SVM-RBF | | | 198 |
| | Logistic Regression | | | 198 |
| | MLP | | | 198 |
| | Gaussian NB | | | 198 |
| | Extra-Trees | | | 198 |
| (synthetic only) | 0.0001 s | | | 144 |
| | 0.001 s | | | 144 |
| | 0.01 s | | | 144 |
| | 0.05 s | | | 144 |
| | 0.1 s | | | 144 |
| | 0.5 s | | | 144 |
| | 1.0 s | | | 144 |
| | 5.0 s | | | 144 |
Table 13.
Method × classifier mean balanced accuracy by obfuscation method and attacker classifier. Each cell reports mean balanced accuracy averaged across all relevant windows and values for that method-classifier combination. Baseline uses six cells per classifier, synthetic methods use 48 cells per classifier, and donor uses 48 donor-labeled cells per classifier, corresponding to six unique donor configurations repeated across labels. The best attacker per row is shown in bold and the worst attacker per row is italicized.
Table 13.
Method × classifier mean balanced accuracy by obfuscation method and attacker classifier. Each cell reports mean balanced accuracy averaged across all relevant windows and values for that method-classifier combination. Baseline uses six cells per classifier, synthetic methods use 48 cells per classifier, and donor uses 48 donor-labeled cells per classifier, corresponding to six unique donor configurations repeated across labels. The best attacker per row is shown in bold and the worst attacker per row is italicized.
| Method | RF | XGBoost | Extra-Trees | KNN | SVM-RBF | LogReg | GNB | MLP |
|---|
| Baseline | 0.850 | 0.855 | 0.852 | 0.804 | 0.825 | 0.814 | 0.841 | 0.737 |
| Fixed | 0.927 | 0.925 | 0.922 | 0.912 | 0.890 | 0.888 | 0.881 | 0.818 |
| Exponential | 0.923 | 0.920 | 0.922 | 0.837 | 0.850 | 0.846 | 0.877 | 0.754 |
| Uniform | 0.924 | 0.925 | 0.925 | 0.858 | 0.861 | 0.851 | 0.878 | 0.757 |
| Donor | 0.093 | 0.090 | 0.108 | 0.130 | 0.168 | 0.170 | 0.294 | 0.307 |
Table 14.
Nash equilibrium strategies ().
Table 14.
Nash equilibrium strategies ().
| Player | Strategy | Probability | Role |
|---|
| Defender | Donor s/ | 0.7245 | Primary defense |
| Defender | Donor s/ | 0.2755 | Secondary defense |
| Attacker | Gaussian NB | 0.7283 | Primary classifier |
| Attacker | MLP | 0.2717 | Secondary classifier |
Table 15.
Attacker training costs.
Table 15.
Attacker training costs.
| Classifier | Family | Raw Time (s) | | In NE? |
|---|
| KNN | Instance-Based | 0.001 | 0.002 | No |
| Gaussian NB | Probabilistic | 0.002 | 0.003 | Yes (72.8%) |
| Logistic Reg. | Linear | 0.070 | 0.110 | No |
| SVM-RBF | Kernel | 0.105 | 0.164 | No |
| MLP | Neural Network | 0.285 | 0.444 | Yes (27.2%) |
| Extra-Trees | Randomized Ens. | 0.359 | 0.559 | No |
| XGBoost | Gradient Boost. | 0.388 | 0.605 | No |
| Random Forest | Classical Ens. | 0.641 | 1.000 | No |
Table 16.
Pairing-aware attacker evaluation across the six unique donor configurations. Each column reports the strongest attacker classifier for that metric. Oracle relabeling applies the balanced-accuracy-maximizing label permutation to the classifier output; pair membership is the two-superclass problem {Bulb, Plug} vs. {Camera, Doorbell}; within-pair is the two-class separation inside each pair after relabeling.
Table 16.
Pairing-aware attacker evaluation across the six unique donor configurations. Each column reports the strongest attacker classifier for that metric. Oracle relabeling applies the balanced-accuracy-maximizing label permutation to the classifier output; pair membership is the two-superclass problem {Bulb, Plug} vs. {Camera, Doorbell}; within-pair is the two-class separation inside each pair after relabeling.
| Window | Best Pairing-Unaware | Best Oracle Relabeling | Pair Membership | Within-Pair (Idle) | Within-Pair (Active) |
|---|
| 1 s | 0.365 (GNB) | 0.888 (RF) | 1.000 | 0.822 | 0.999 |
| 5 s | 0.373 (GNB) | 0.852 (SVM-RBF) | 0.999 | 0.824 | 0.999 |
| 10 s | 0.335 (MLP) | 0.865 (SVM-RBF) | 0.997 | 0.860 | 0.995 |
| 30 s | 0.460 (MLP) | 0.959 (KNN) | 0.996 | 0.949 | 1.000 |
| 60 s | 0.264 (GNB) | 0.975 (XGBoost) | 0.985 | 1.000 | 1.000 |
| 120 s | 0.401 (MLP) | 0.915 (Extra-Trees) | 0.972 | 1.000 | 0.944 |