Figure 1.
Proposed DL-based framework DeepMedShield-XAI with PSO.
Figure 1.
Proposed DL-based framework DeepMedShield-XAI with PSO.
Figure 2.
Flow chart of the proposed framework DeepMedShield-XAI.
Figure 2.
Flow chart of the proposed framework DeepMedShield-XAI.
Figure 3.
Number of samples in 6 categories of the CICIoMT2024 dataset.
Figure 3.
Number of samples in 6 categories of the CICIoMT2024 dataset.
Figure 4.
Number of samples in 9 categories of the IoMT_TrafficData dataset.
Figure 4.
Number of samples in 9 categories of the IoMT_TrafficData dataset.
Figure 5.
Proposed DNN architecture.
Figure 5.
Proposed DNN architecture.
Figure 6.
Proposed CNN architecture.
Figure 6.
Proposed CNN architecture.
Figure 7.
(a) Aggregated accuracy and (b) aggregated loss comparison for the DNN model on the CICIoMT2024 dataset.
Figure 7.
(a) Aggregated accuracy and (b) aggregated loss comparison for the DNN model on the CICIoMT2024 dataset.
Figure 8.
Comparison of (a) ROC curves and (b) PR curves for the DNN model on the CICIoMT2024 dataset.
Figure 8.
Comparison of (a) ROC curves and (b) PR curves for the DNN model on the CICIoMT2024 dataset.
Figure 9.
(a) Aggregated accuracy and (b) aggregated loss comparison for the CNN model on the CICIoMT2024 dataset.
Figure 9.
(a) Aggregated accuracy and (b) aggregated loss comparison for the CNN model on the CICIoMT2024 dataset.
Figure 10.
Comparison of (a) ROC curves and (b) PR curves for the CNN model on the CICIoMT2024 dataset.
Figure 10.
Comparison of (a) ROC curves and (b) PR curves for the CNN model on the CICIoMT2024 dataset.
Figure 11.
(a) Aggregated accuracy and (b) aggregated loss comparison for the encoder–transformer model on the CICIoMT2024 dataset.
Figure 11.
(a) Aggregated accuracy and (b) aggregated loss comparison for the encoder–transformer model on the CICIoMT2024 dataset.
Figure 12.
Comparison of (a) ROC curves and (b) PR curves for the encoder–transformer model on the CICIoMT2024 dataset.
Figure 12.
Comparison of (a) ROC curves and (b) PR curves for the encoder–transformer model on the CICIoMT2024 dataset.
Figure 13.
LIME results of the CICIoMT2024 dataset for the DNN model for sample 0.
Figure 13.
LIME results of the CICIoMT2024 dataset for the DNN model for sample 0.
Figure 14.
LIME results of the CICIoMT2024 dataset for the DNN model for sample 1.
Figure 14.
LIME results of the CICIoMT2024 dataset for the DNN model for sample 1.
Figure 15.
LIME results of the CICIoMT2024 dataset for the DNN model for sample 2.
Figure 15.
LIME results of the CICIoMT2024 dataset for the DNN model for sample 2.
Figure 16.
SHAP explanations for the DNN model: (a) benign class and (b) DDoS class.
Figure 16.
SHAP explanations for the DNN model: (a) benign class and (b) DDoS class.
Figure 17.
SHAP explanations for the DNN model: (a) DoS class and (b) MQTT class.
Figure 17.
SHAP explanations for the DNN model: (a) DoS class and (b) MQTT class.
Figure 18.
(a) SHAP analysis for the recon class and (b) SHAP analysis for the spoofing class for the DNN model.
Figure 18.
(a) SHAP analysis for the recon class and (b) SHAP analysis for the spoofing class for the DNN model.
Figure 19.
SHAP feature importance for the benign class.
Figure 19.
SHAP feature importance for the benign class.
Figure 20.
SHAP feature importance for the DDoS class.
Figure 20.
SHAP feature importance for the DDoS class.
Figure 21.
SHAP feature importance for the DoS class.
Figure 21.
SHAP feature importance for the DoS class.
Figure 22.
SHAP feature importance for the MQTT class.
Figure 22.
SHAP feature importance for the MQTT class.
Figure 23.
SHAP feature importance for the recon class.
Figure 23.
SHAP feature importance for the recon class.
Figure 24.
SHAP feature importance for the spoofing class.
Figure 24.
SHAP feature importance for the spoofing class.
Figure 25.
(a) Aggregated accuracy and (b) aggregated loss comparison for the DNN model on the IoMT_TrafficData dataset.
Figure 25.
(a) Aggregated accuracy and (b) aggregated loss comparison for the DNN model on the IoMT_TrafficData dataset.
Figure 26.
Comparison of (a) ROC curves and (b) PR curves for the DNN model on the IoMT_TrafficData dataset.
Figure 26.
Comparison of (a) ROC curves and (b) PR curves for the DNN model on the IoMT_TrafficData dataset.
Figure 27.
(a) Aggregated accuracy and (b) aggregated loss comparison for the CNN model on the IoMT_TrafficData dataset.
Figure 27.
(a) Aggregated accuracy and (b) aggregated loss comparison for the CNN model on the IoMT_TrafficData dataset.
Figure 28.
Comparison of (a) ROC curves and (b) PR curves for the CNN model on the IoMT_TrafficData dataset.
Figure 28.
Comparison of (a) ROC curves and (b) PR curves for the CNN model on the IoMT_TrafficData dataset.
Figure 29.
(a) Aggregated accuracy and (b) aggregated loss comparison for the encoder–transformer model on the IoMT_TrafficData dataset.
Figure 29.
(a) Aggregated accuracy and (b) aggregated loss comparison for the encoder–transformer model on the IoMT_TrafficData dataset.
Figure 30.
Comparison of (a) ROC curves and (b) PR curves for the encoder–transformer model on the IoMT_TrafficData dataset.
Figure 30.
Comparison of (a) ROC curves and (b) PR curves for the encoder–transformer model on the IoMT_TrafficData dataset.
Figure 31.
LIME results of IoMT_TrafficData dataset for the DNN model for sample 0.
Figure 31.
LIME results of IoMT_TrafficData dataset for the DNN model for sample 0.
Figure 32.
LIME results of IoMT_TrafficData dataset for the DNN model for sample 1.
Figure 32.
LIME results of IoMT_TrafficData dataset for the DNN model for sample 1.
Figure 33.
LIME results of IoMT_TrafficData dataset for the DNN model for sample 2.
Figure 33.
LIME results of IoMT_TrafficData dataset for the DNN model for sample 2.
Figure 34.
(a) SHAP analysis for the apachekiller class and (b) SHAP for the arpspoofing class for the DNN model.
Figure 34.
(a) SHAP analysis for the apachekiller class and (b) SHAP for the arpspoofing class for the DNN model.
Figure 35.
(a) SHAP analysis for the camoverflow class and (b) SHAP analysis for the mqttmalaria class for the DNN model.
Figure 35.
(a) SHAP analysis for the camoverflow class and (b) SHAP analysis for the mqttmalaria class for the DNN model.
Figure 36.
(a) SHAP analysis for the netscan class and (b) SHAP analysis for the normal class for the DNN model.
Figure 36.
(a) SHAP analysis for the netscan class and (b) SHAP analysis for the normal class for the DNN model.
Figure 37.
(a) SHAP analysis for the rudeadyet class and (b) SHAP for the slowloris class for the DNN model.
Figure 37.
(a) SHAP analysis for the rudeadyet class and (b) SHAP for the slowloris class for the DNN model.
Figure 38.
SHAP analysis for the slowread class for the DNN model.
Figure 38.
SHAP analysis for the slowread class for the DNN model.
Figure 39.
SHAP feature importance for the apachekiller class.
Figure 39.
SHAP feature importance for the apachekiller class.
Figure 40.
SHAP feature importance for the arpspoofing class.
Figure 40.
SHAP feature importance for the arpspoofing class.
Figure 41.
SHAP feature importance for the camoverflow class.
Figure 41.
SHAP feature importance for the camoverflow class.
Figure 42.
SHAP feature importance for the mqttmalaria class.
Figure 42.
SHAP feature importance for the mqttmalaria class.
Figure 43.
SHAP feature importance for the netscan class.
Figure 43.
SHAP feature importance for the netscan class.
Figure 44.
SHAP feature importance for the rudeadyet class.
Figure 44.
SHAP feature importance for the rudeadyet class.
Figure 45.
SHAP feature importance for the slowloris class.
Figure 45.
SHAP feature importance for the slowloris class.
Figure 46.
SHAP feature importance for the slowread class.
Figure 46.
SHAP feature importance for the slowread class.
Figure 47.
SHAP feature importance for the normal class.
Figure 47.
SHAP feature importance for the normal class.
Table 1.
Five-fold cross-validation results with training time.
Table 1.
Five-fold cross-validation results with training time.
| Dataset | Model | Fold | Accuracy (%) | Precision (%) | Recall (%) | F1-Score (%) | Epochs | Time (m) |
|---|
| CICIoMT | DNN | 1 | 99.63 | 91.41 | 89.86 | 90.56 | 99 | 74.25 |
| 2 | 99.61 | 91.15 | 89.66 | 90.34 | 90 | 69 |
| 3 | 99.71 | 93.40 | 93.10 | 93.24 | 150 | 120 |
| 4 | 99.69 | 92.32 | 93.02 | 92.63 | 166 | 132.8 |
| 5 | 99.64 | 91.66 | 90.36 | 90.96 | 146 | 116.8 |
| Mean ± Std | 99.65 ± 0.04 | 91.99 ± 0.90 | 91.20 ± 1.72 | 91.55 ± 1.30 | - | 512.85 |
| CNN | 1 | 78.94 | 77.52 | 83.02 | 79.33 | 23 | 22.61 |
| 2 | 78.78 | 78.02 | 80.99 | 77.84 | 21 | 21 |
| 3 | 76.23 | 76.81 | 78.31 | 73.23 | 22 | 22 |
| 4 | 75.30 | 74.73 | 74.77 | 66.87 | 21 | 21 |
| 5 | 74.65 | 75.82 | 76.83 | 70.55 | 22 | 22 |
| Mean ± Std | 76.78 ± 1.98 | 76.58 ± 1.32 | 78.79 ± 3.28 | 73.56 ± 5.14 | - | 108.61 |
| ET | 1 | 99.12 | 90.40 | 88.76 | 89.49 | 62 | 84.73 |
| 2 | 99.29 | 90.44 | 88.47 | 89.33 | 29 | 39.63 |
| 3 | 99.44 | 90.76 | 88.73 | 89.62 | 65 | 88.83 |
| 4 | 99.43 | 90.56 | 88.49 | 89.38 | 23 | 31.05 |
| 5 | 99.50 | 90.91 | 90.99 | 90.91 | 39 | 55.9 |
| Mean ± Std | 99.36 ± 0.15 | 90.61 ± 0.22 | 89.09 ± 1.07 | 89.75 ± 0.66 | - | 300.14 |
| IoMT | DNN | 1 | 99.89 | 98.97 | 99.20 | 99.09 | 47 | 16.45 |
| 2 | 99.87 | 98.59 | 99.26 | 98.92 | 30 | 10.5 |
| 3 | 99.86 | 98.90 | 99.03 | 98.96 | 43 | 15.05 |
| 4 | 99.87 | 98.65 | 99.21 | 98.92 | 33 | 11.05 |
| 5 | 99.88 | 99.17 | 98.98 | 99.07 | 22 | 7.7 |
| Mean ± Std | 99.87 ± 0.01 | 98.86 ± 0.24 | 99.14 ± 0.12 | 98.99 ± 0.08 | - | 60.75 |
| CNN | 1 | 99.88 | 99.22 | 98.85 | 99.03 | 38 | 17.1 |
| 2 | 99.87 | 99.26 | 98.98 | 99.12 | 12 | 5.6 |
| 3 | 99.87 | 99.00 | 99.02 | 99.01 | 16 | 7.4 |
| 4 | 99.88 | 99.24 | 98.97 | 99.10 | 13 | 5.85 |
| 5 | 99.88 | 99.06 | 99.07 | 99.06 | 21 | 9.45 |
| Mean ± Std | 99.88 ± 0.01 | 99.16 ± 0.12 | 98.98 ± 0.08 | 99.06 ± 0.05 | - | 45.4 |
| ET | 1 | 99.80 | 99.00 | 98.09 | 98.54 | 100 | 141.67 |
| 2 | 99.81 | 98.93 | 98.47 | 98.70 | 31 | 43.92 |
| 3 | 99.79 | 98.97 | 98.02 | 98.49 | 62 | 87.83 |
| 4 | 99.83 | 98.91 | 98.52 | 98.71 | 82 | 116.17 |
| 5 | 99.83 | 98.70 | 98.51 | 98.59 | 53 | 72.25 |
| Mean ± Std | 99.81 ± 0.02 | 98.90 ± 0.12 | 98.32 ± 0.25 | 98.61 ± 0.10 | - | 461.84 |
Table 2.
Performance analysis of DNN, CNN, and ET models on CICIoMT and IoMT datasets.
Table 2.
Performance analysis of DNN, CNN, and ET models on CICIoMT and IoMT datasets.
| Dataset | Model | Fold | Bal_Acc (%) | MCC (%) | Kappa (%) | G-Mean (%) | Jaccard (%) | Dice (%) |
|---|
| CICIoMT | DNN | 1 | 89.87 | 99.24 | 99.24 | 94.76 | 86.27 | 90.57 |
| 2 | 89.66 | 99.18 | 99.18 | 94.65 | 86.01 | 90.34 |
| 3 | 93.08 | 99.41 | 99.41 | 96.45 | 89.20 | 93.12 |
| 4 | 93.02 | 99.36 | 99.36 | 96.41 | 88.60 | 92.62 |
| 5 | 90.36 | 99.26 | 99.26 | 95.02 | 86.66 | 90.96 |
| Mean ± Std | 91.20 ± 1.71 | 99.29 ± 0.09 | 99.29 ± 0.09 | 95.46 ± 0.90 | 87.35 ± 1.45 | 91.52 ± 1.26 |
| 95% CI | [89.07, 93.32] | [99.17, 99.41] | [99.17, 99.41] | [94.34, 96.57] | [85.55, 89.15] | [89.95, 93.09] |
| CNN | 1 | 83.02 | 56.92 | 56.92 | 87.62 | 69.47 | 79.33 |
| 2 | 80.96 | 54.07 | 52.51 | 85.85 | 68.28 | 77.81 |
| 3 | 78.33 | 48.49 | 42.37 | 83.47 | 64.42 | 73.22 |
| 4 | 73.21 | 40.49 | 28.83 | 79.76 | 55.62 | 63.43 |
| 5 | 76.83 | 45.66 | 35.74 | 82.12 | 62.85 | 70.55 |
| Mean ± Std | 78.47 ± 3.78 | 49.13 ± 6.56 | 43.27 ± 11.59 | 83.76 ± 3.08 | 64.13 ± 5.47 | 72.87 ± 6.34 |
| 95% CI | [73.77, 83.17] | [40.98, 57.27] | [28.88, 57.67] | [79.93, 87.59] | [57.33, 70.93] | [65.00, 80.74] |
| ET | 1 | 88.75 | 98.19 | 98.19 | 94.07 | 84.75 | 89.49 |
| 2 | 88.41 | 98.32 | 98.32 | 93.91 | 84.56 | 89.28 |
| 3 | 88.73 | 98.85 | 98.85 | 94.14 | 85.02 | 89.62 |
| 4 | 88.49 | 98.83 | 98.83 | 94.01 | 84.86 | 89.38 |
| 5 | 90.99 | 98.98 | 98.98 | 95.33 | 86.47 | 90.91 |
| Mean ± Std | 89.07 ± 1.08 | 98.64 ± 0.35 | 98.63 ± 0.35 | 94.29 ± 0.59 | 85.13 ± 0.77 | 89.74 ± 0.67 |
| 95% CI | [87.73, 90.42] | [98.20, 99.07] | [98.19, 99.07] | [93.56, 95.02] | [84.18, 86.08] | [88.91, 90.57] |
| IoMT | DNN | 1 | 88.25 | 88.70 | 87.72 | 93.16 | 86.04 | 87.49 |
| 2 | 99.25 | 99.65 | 99.65 | 99.61 | 97.98 | 98.97 |
| 3 | 99.01 | 99.60 | 99.60 | 99.48 | 97.95 | 98.95 |
| 4 | 99.18 | 99.63 | 99.63 | 99.57 | 97.86 | 98.90 |
| 5 | 98.96 | 99.64 | 99.64 | 99.46 | 98.15 | 99.05 |
| Mean ± Std | 96.93 ± 4.85 | 97.44 ± 4.89 | 97.25 ± 5.32 | 98.26 ± 2.85 | 95.60 ± 5.34 | 96.67 ± 5.13 |
| 95% CI | [90.91, 100] | [91.37, 100] | [90.64, 100] | [94.72, 100] | [88.96, 100] | [90.30, 100] |
| CNN | 1 | 98.83 | 99.66 | 99.66 | 99.40 | 98.06 | 99.01 |
| 2 | 98.95 | 99.64 | 99.64 | 99.46 | 98.24 | 99.10 |
| 3 | 99.00 | 99.63 | 99.63 | 99.48 | 98.02 | 98.99 |
| 4 | 98.94 | 99.66 | 99.66 | 99.45 | 98.22 | 99.09 |
| 5 | 99.05 | 99.67 | 99.67 | 99.51 | 98.13 | 99.04 |
| Mean ± Std | 98.95 ± 0.08 | 99.65 ± 0.02 | 99.65 ± 0.02 | 99.46 ± 0.04 | 98.13 ± 0.09 | 99.05 ± 0.05 |
| 95% CI | [98.86, 99.05] | [99.63, 99.67] | [99.63, 99.67] | [99.41, 99.51] | [98.02, 98.25] | [98.99, 99.11] |
| ET | 1 | 98.06 | 99.42 | 99.41 | 99.00 | 97.15 | 98.53 |
| 2 | 98.45 | 99.45 | 99.44 | 99.19 | 97.44 | 98.68 |
| 3 | 97.99 | 99.40 | 99.40 | 98.96 | 97.05 | 98.47 |
| 4 | 98.50 | 99.52 | 99.52 | 99.22 | 97.47 | 98.69 |
| 5 | 98.49 | 99.52 | 99.52 | 99.22 | 97.26 | 98.58 |
| Mean ± Std | 98.30 ± 0.25 | 99.46 ± 0.06 | 99.46 ± 0.06 | 99.12 ± 0.13 | 97.27 ± 0.18 | 98.59 ± 0.10 |
| 95% CI | [97.99, 98.61] | [99.39, 99.53] | [99.39, 99.53] | [98.95, 99.28] | [97.05, 97.50] | [98.47, 98.71] |
Table 3.
Test performance comparison of DNN, CNN, and ET models (in %).
Table 3.
Test performance comparison of DNN, CNN, and ET models (in %).
| Dataset | Model | Accuracy (%) | Precision (%) | Recall (%) | F1-Score (%) | ROC-AUC |
|---|
| CICIoMT2024 | DNN | 99.68 | 91.99 | 91.15 | 91.54 | 0.9999 |
| CICIoMT2024 | CNN | 76.78 | 76.58 | 78.79 | 73.5 | 0.9592 |
| CICIoMT2024 | ET | 99.35 | 90.56 | 89.01 | 89.72 | 0.9981 |
| IoMT | DNN | 99.87 | 99.05 | 98.95 | 98.99 | 0.9998 |
| IoMT | CNN | 99.87 | 98.97 | 99.04 | 99.00 | 0.9999 |
| IoMT | ET | 99.82 | 98.72 | 98.51 | 98.60 | 0.9984 |
Table 4.
Robustness analysis of the DNN model under FGSM and PGD adversarial attacks on CICIoMT and IoMT datasets.
Table 4.
Robustness analysis of the DNN model under FGSM and PGD adversarial attacks on CICIoMT and IoMT datasets.
| Dataset | Fold | FGSM Acc. (%) | FGSM Drop (%) | PGD Acc. (%) | PGD Drop (%) | OOD Rate (%) |
|---|
| DNN_IoMT | 1 | 87.15 | 12.72 | 99.75 | 0.12 | 0.28 |
| 2 | 91.02 | 8.86 | 99.45 | 0.43 | 0.29 |
| 3 | 95.35 | 4.52 | 99.72 | 0.15 | 0.33 |
| 4 | 92.92 | 6.96 | 99.55 | 0.33 | 0.32 |
| 5 | 95.69 | 4.19 | 99.59 | 0.29 | 0.32 |
| Mean ± Std | 92.43 ± 3.14 | 7.45 ± 3.14 | 99.61 ± 0.11 | 0.26 ± 0.11 | 0.31 ± 0.02 |
| 95% CI | [88.53, 96.33] | [3.55, 11.35] | [99.47, 99.75] | [0.12, 0.40] | [0.28, 0.34] |
| DNN_CICIoMT | 1 | 99.52 | 0.09 | 99.52 | 0.08 | 0.95 |
| 2 | 99.42 | 0.15 | 99.43 | 0.14 | 0.96 |
| 3 | 99.46 | 0.14 | 99.47 | 0.13 | 0.89 |
| 4 | 99.51 | 0.08 | 99.51 | 0.08 | 0.94 |
| 5 | 99.36 | 0.31 | 99.36 | 0.31 | 0.93 |
| Mean ± Std | 99.45 ± 0.06 | 0.15 ± 0.08 | 99.46 ± 0.06 | 0.15 ± 0.09 | 0.94 ± 0.02 |
| 95% CI | [99.38, 99.52] | [0.05, 0.25] | [99.39, 99.53] | [0.04, 0.26] | [0.91, 0.97] |
Table 5.
Statistical significance analysis using balanced accuracy and F1-score across five-fold cross-validation results.
Table 5.
Statistical significance analysis using balanced accuracy and F1-score across five-fold cross-validation results.
| Dataset | Metric | Comparison | Shapiro-p | t-Stat | p-Value | Wilcoxon Stat | Wilcoxon-p |
|---|
| CICIoMT | Balanced Accuracy | DNN vs. CNN | 0.8112 | 5.5375 | 0.0052 | 0 | 0.0625 |
| CICIoMT | Balanced Accuracy | DNN vs. ET | 0.3158 | 2.1187 | 0.1015 | 1 | 0.1250 |
| CICIoMT | Balanced Accuracy | CNN vs. ET | 0.5957 | −5.7401 | 0.0046 | 0 | 0.0625 |
| CICIoMT | F1-score | DNN vs. CNN | 0.5257 | 5.7828 | 0.0044 | 0 | 0.0625 |
| CICIoMT | F1-score | DNN vs. ET | 0.2840 | 2.6490 | 0.0570 | 0 | 0.0625 |
| CICIoMT | F1-score | CNN vs. ET | 0.7033 | −5.8022 | 0.0044 | 0 | 0.0625 |
| IoMT | Balanced Accuracy | DNN vs. CNN | 0.0004 | −0.9457 | 0.3978 | 7 | 1.0000 |
| IoMT | Balanced Accuracy | DNN vs. ET | 0.0005 | −0.6476 | 0.5525 | 5 | 0.6250 |
| IoMT | Balanced Accuracy | CNN vs. ET | 0.4221 | 6.2759 | 0.0033 | 0 | 0.0625 |
| IoMT | F1-score | DNN vs. CNN | 0.0002 | −1.0381 | 0.3578 | 1 | 0.1250 |
| IoMT | F1-score | DNN vs. ET | 0.0003 | −0.8408 | 0.4478 | 5 | 0.6250 |
| IoMT | F1-score | CNN vs. ET | 0.8989 | 21.3542 | 2.84 × 10−5 | 0 | 0.0625 |
Table 6.
Performance comparison of existing methods and the proposed model on IoMT datasets.
Table 6.
Performance comparison of existing methods and the proposed model on IoMT datasets.
| Research | Dataset | Method | Accuracy (%) | AUC |
|---|
| Hafid et al. [14] | CICIoMT2024 | XGBoost + SHAP | 97 | – |
| Torre et al. [23] | CICIoMT2024 | CNN | 97.31 | – |
| Rehaman et al. [24] | CICIoMT2024 | IG, MI, Fisher Score + ML | 97.7 | – |
| | IoMT_TrafficData | | 98.7 | – |
| Akar et al. [25] | CICIoMT2024 | LSTM | 98 | – |
| Bo et al. [26] | CICIoMT2024 | Meta-learning + Feature Fusion | 97.78 | – |
| Palaniappan et al. [27] | CICIoMT2024 | ProxyNet + MI Filtering + ML | – | 0.9998 |
| | IoMT_TrafficData | | – | 0.9997 |
| Proposed | CICIoMT2024 | PSO + Deep Learning + XAI (5-fold CV) | | |
| | IoMT | | | |