Evaluating Model Resilience to Data Poisoning Attacks: A Comparative Study
Abstract
1. Introduction
2. Related Work
2.1. Data Poisoning Attacks in Machine Learning
2.2. Defenses and Robust Training Strategies
2.3. Explainable AI for Robustness and Security
2.4. Theoretical Impact on Learning Dynamics
2.5. Broader Context in Machine Learning System Security
3. Methodology
3.1. Overview of the Proposed Framework
3.2. Preprocessing & Reproducibility Details
3.3. Model Definitions and Implementation Framework
3.3.1. Logistic Regression
3.3.2. Decision Tree
3.3.3. Vanilla Multi-Layer Perceptron (MLP)
3.3.4. Regularized MLP
3.3.5. Convolutional Neural Network (CNN)
3.3.6. Bayesian Neural Network with LSTM (BNN+LSTM)
3.3.7. DistilBERT with LoRA Fine-Tuning
3.4. Poisoning Injection Algorithms
3.4.1. Label Flipping
3.4.2. Data Corruption
3.4.3. Adversarial Insertion
3.4.4. Loss Functions
3.5. Explainability Algorithm for Failure Analysis
| Algorithm 1: Cross-architectural evaluation of resilience to data poisoning |
![]() |
4. Evaluation Result
4.1. Dataset and Settings
4.2. Model Initialization
4.2.1. Classifier Models
4.2.2. Explainability Model
4.3. Evaluation Metrics
Quantitative Metrics
4.4. Performance Evaluation Under Varying Poisoning Severity Levels
4.5. Comparative Robustness Under Selected Poisoning Level
4.6. Mechanistic Interpretation of Degradation Patterns
4.7. Cross-Model Behavior Summary
4.8. Case Study with LIME Explanations
5. Discussion
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Tavallali, P.; Behzadan, V.; Alizadeh, A.; Ranganath, A.; Singhal, M. Adversarial Label-Poisoning Attacks and Defense for General Multi-Class Models Based On Synthetic Reduced Nearest Neighbor. In Proceedings of the International Conference on Image Processing, ICIP, Bordeaux, France, 16–19 October 2022; IEEE Computer Society: New York, NY, USA; pp. 3717–3722. [Google Scholar] [CrossRef] [Scilit]
- Lu, Y.; Kamath, G.; Yu, Y. Exploring the Limits of Model-Targeted Indiscriminate Data Poisoning Attacks. arXiv 2023, arXiv:2303.03592. [Google Scholar]
- Zhang, H.; Cheng, N.; Zhang, Y.; Li, Z. Label flipping attacks against Naive Bayes on spam filtering systems. Appl. Intell. 2021, 51, 4503–4514. [Google Scholar] [CrossRef] [Scilit]
- Manthena, H.; Shajarian, S.; Kimmell, J.; Abdelsalam, M.; Khorsandroo, S.; Gupta, M. Explainable Artificial Intelligence (XAI) for Malware Analysis: A Survey of Techniques, Applications, and Open Challenges. arXiv 2024, arXiv:2409.13723. [Google Scholar] [CrossRef] [Scilit]
- Mengara, O. A backdoor approach with inverted labels using dirty label-flipping attacks. IEEE Access 2024, 13, 124225–124233. [Google Scholar] [CrossRef] [Scilit]
- Ji, J. Investigating the Label-flipping Attacks Impact in Federated Learning. In Proceedings of the 2024 5th International Conference on Information Science, Parallel and Distributed Systems, ISPDS, Guangzhou, China, 31 May–2 June 2024; Institute of Electrical and Electronics Engineers: New York, NY, USA 2024; pp. 82–86. [Google Scholar] [CrossRef] [Scilit]
- Truong, L.; Jones, C.; Hutchinson, B.; August, A.; Praggastis, B.; Jasper, R.; Nichols, N.; Tuor, A. Systematic Evaluation of Backdoor Data Poisoning Attacks on Image Classifiers. arXiv 2020, arXiv:2004.11514. [Google Scholar] [CrossRef] [Scilit]
- Xue, J.; Zheng, M.; Hua, T.; Shen, Y.; Liu, Y.; Boloni, L.; Lou, Q. TrojLLM: A Black-box Trojan Prompt Attack on Large Language Models. arXiv 2023, arXiv:2306.06815. [Google Scholar]
- Jebreel, M.; Mukkamala, R.R.; Vatrapu, R. Defending Label-Flipping Attacks in Federated Learning. In Proceedings of the 2022 IEEE International Conference on Big Data (Big Data), Osaka, Japan, 17–20 December 2022; IEEE: New York, NY, USA; pp. 4320–4327. [Google Scholar] [CrossRef] [Scilit]
- Liu, T.Y.; Yang, Y.; Mirzasoleiman, B. Friendly Noise against Adversarial Noise: A Powerful Defense against Data Poisoning Attacks. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA, 28 November–9 December 2022. [Google Scholar]
- Mahbooba, B.; Timilsina, M.; Sahal, R.; Serrano, M. Explainable artificial intelligence (XAI) to enhance trust management in intrusion detection systems using decision tree model. Complexity 2021, 2021, 6634811. [Google Scholar] [CrossRef] [Scilit]
- Ali, M.; Zhang, J. Exploring the Effectiveness of Synthetic Data in Network Intrusion Detection through XAI. In Proceedings of the 2024 Cyber Awareness and Research Symposium (CARS), Grand Forks, ND, USA, 28–29 October 2024; pp. 1–5. [Google Scholar]
- Patil, S.; Varadarajan, V.; Mazhar, S.M.; Sahibzada, A.; Ahmed, N.; Sinha, O.; Kumar, S.; Shaw, K.; Kotecha, K. Explainable artificial intelligence for intrusion detection system. Electronics 2022, 11, 3079. [Google Scholar] [CrossRef] [Scilit]
- Arreche, O.; Guntur, T.; Abdallah, M. Xai-ids: Toward proposing an explainable artificial intelligence framework for enhancing network intrusion detection systems. Appl. Sci. 2024, 14, 4170. [Google Scholar] [CrossRef] [Scilit]
- Gummadi, A.N.; Napier, J.C.; Abdallah, M. XAI-IoT: An explainable AI framework for enhancing anomaly detection in IoT systems. IEEE Access 2024, 12, 71024–71054. [Google Scholar] [CrossRef] [Scilit]
- Nazat, S.; Li, L.; Abdallah, M. XAI-ADS: An explainable artificial intelligence framework for enhancing anomaly detection in autonomous driving systems. IEEE Access 2024, 12, 48583–48607. [Google Scholar] [CrossRef] [Scilit]
- Dunn, C.; Moustafa, N.; Turnbull, B. Robustness evaluations of sustainable machine learning models against data poisoning attacks in the internet of things. Sustainability 2020, 12, 6434. [Google Scholar] [CrossRef] [Scilit]
- Anisetti, M.; Ardagna, C.A.; Balestrucci, A.; Bena, N.; Damiani, E.; Yeun, C.Y. On the Robustness of Random Forest Against Untargeted Data Poisoning: An Ensemble-Based Approach. IEEE Trans. Sustain. Comput. 2023, 8, 540–554. [Google Scholar] [CrossRef] [Scilit]
- Insua, D.R.; Naveiro, R.; Gallego, V.; Poulos, J. Adversarial machine learning: Bayesian perspectives. arXiv 2020, arXiv:2003.03546. [Google Scholar]
- Pawelczyk, M.; Di, J.Z.; Lu, Y.; Kamath, G.; Sekhari, A.; Neel, S. Machine Unlearning Fails to Remove Data Poisoning Attacks. arXiv 2024, arXiv:2406.17216. [Google Scholar] [CrossRef] [Scilit]
- Familoni, B.T. Cybersecurity challenges in the age of AI: Theoretical approaches and practical solutions. Comput. Sci. Res. J. 2024, 5, 703–724. [Google Scholar]
- Cinà, A.E.; Grosse, K.; Demontis, A.; Biggio, B.; Roli, F.; Pelillo, M. Machine Learning Security Against Data Poisoning: Are We There Yet? Computer 2024, 57, 26–34. [Google Scholar] [CrossRef] [Scilit]
- Alahmed, S.; Alasad, Q.; Yuan, J.S.; Alawad, M. Impacting Robustness in Deep Learning-Based NIDS through Poisoning Attacks. Algorithms 2024, 17, 155. [Google Scholar] [CrossRef] [Scilit]
- Cheng, H.; Fan, Y.; Wang, Z.; Guo, Y.; Wu, J.; Jiang, J.; Zhang, X. Semi-supervised learning with reweighting for robust deep learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Virtually, 2–9 February 2021; Volume 35, pp. 6925–6933. [Google Scholar]
- Yazdinejad, A.; Dehghantanha, A.; Karimipour, H.; Srivastava, G.; Parizi, R.M. A robust privacy-preserving federated learning model against model poisoning attacks. IEEE Trans. Inf. Forensics Secur. 2024, 19, 6693–6708. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Zhao, J.; LeCun, Y. Character-level convolutional networks for text classification. Adv. Neural Inf. Process. Syst. 2015, 28, 649–657. [Google Scholar]
- Zhou, C.; Zhang, M.; Li, J.; Liu, Y.; Chen, T.; Liu, T. LoRA: Low-Rank Adaptation of Large Language Models. arXiv 2023, arXiv:2106.09685. [Google Scholar]
- Liu, X.; Si, S.; Zhu, X.; Li, Y.; Hsieh, C.J. A Unified Framework for Data Poisoning Attack to Graph-based Semi-supervised Learning. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 8–14 December 2019; pp. 9777–9787. [Google Scholar]
- Sokol, K.; Hepburn, A.; Santos-Rodríguez, R.; Flach, P.A. bLIMEy: Surrogate Prediction Explanations Beyond LIME. arXiv 2019, arXiv:1910.13016. [Google Scholar] [CrossRef] [Scilit]
- Maas, A.L.; Daly, R.E.; Pham, P.T.; Huang, D.; Ng, A.Y.; Potts, C. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Portland, Oregon, 19–24 June 2011; pp. 142–150. [Google Scholar]
- Shovon, A.R.; Sun, Y.; Micinski, K.; Gilray, T.; Kumar, S. Multi-node multi-gpu datalog. In Proceedings of the 39th ACM International Conference on Supercomputing, Salt Lake City, UT, USA, 9–11 June 2025; pp. 822–836. [Google Scholar]
- Moritz, P.; Nishihara, R.; Wang, S.; Tumanov, A.; Liaw, R.; Liang, E.; Elibol, M.; Yang, Z.; Paul, W.; Jordan, M.I.; et al. Ray: A distributed framework for emerging AI applications. In Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18), Carlsbad, CA, USA, 8–10 October 2018; pp. 561–577. [Google Scholar]
- Rezaei, H.; Taheri, R.; Shojafar, M. FedLLMGuard: A federated large language model for anomaly detection in 5G networks. Comput. Netw. 2025, 269, 111473. [Google Scholar] [CrossRef] [Scilit]







| Model Type | Architecture Summary | Training Configuration | Input Format |
|---|---|---|---|
| Logistic Regression | Single-layer linear classifier (L2 regularization) via scikit-learn | Full-batch, 1000 iterations, liblinear solver | TF-IDF vector (5k features) |
| Decision Tree | Unbounded-depth tree using Gini impurity | Default hyperparameters | TF-IDF vector (5k features) |
| Vanilla MLP | 3-layer FC network with ReLU and softmax | 50 epochs, Adam (), batch size 64 | Tokenized padded sequences |
| Regularized MLP | Same as Vanilla MLP + Dropout () and L2 regularization | Same as above | Tokenized padded sequences |
| CNN | Embedding → Conv1D (128) → MaxPool → FC | 15 epochs, batch size 64, input length 300 tokens | Tokenized padded sequences |
| BNN+LSTM | Embedding → Bayesian LSTM → Dropout → BayesianLinear | 65 epochs, Adam, KL-regularized loss () | Tokenized padded sequences |
| DistilBERT + LoRA | DistilBERT fine-tuned using LoRA attention injection | 30 epochs, batch size 32, using Trainer API | Tokenized inputs via Hugging Face |
| Model Type | Wrapper or Interface Used | LIME Explanation Notes |
|---|---|---|
| Logistic Regression | make_pipeline(vectorizer, model) | Uses predict_proba() for token attribution |
| Vanilla/Regularized MLP | Custom PyTorch wrapper returning softmax(model(x)) | Evaluated using LIME post Phase 3 |
| CNN | TorchModelWrapper with padding + tokenization | Returns softmax scores, supports all attack types |
| BNN+LSTM | TorchBNNLSTMWrapper with stochastic forward passes | Uses posterior mean for softmax-based attribution |
| Model Group | Most Vulnerable To | Most Resilient Against | Behavioral Observation |
|---|---|---|---|
| LogReg, Decision Tree | Label Flipping | Data Corruption | Overfocus on syntactic or frequent neutral tokens |
| Vanilla MLP | Adversarial Insertion | Data Corruption | Diffused attention, unstable under semantic conflict |
| Regularized MLP | Label Flipping | Data Corruption | Regularization stabilizes training and limits drift |
| CNN | Label Flipping | Adversarial Insertion | Filter collapse under polarity confusion |
| BNN+LSTM | None | All types | Probabilistic attribution shift improves resilience |
| LLM (DistilBERT) | None | All types | Attention reweighting supports context-aware predictions |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Udoidiok, I.; Li, F.; Zhang, J. Evaluating Model Resilience to Data Poisoning Attacks: A Comparative Study. Information 2026, 17, 9. https://doi.org/10.3390/info17010009
Udoidiok I, Li F, Zhang J. Evaluating Model Resilience to Data Poisoning Attacks: A Comparative Study. Information. 2026; 17(1):9. https://doi.org/10.3390/info17010009
Chicago/Turabian StyleUdoidiok, Ifiok, Fuhao Li, and Jielun Zhang. 2026. "Evaluating Model Resilience to Data Poisoning Attacks: A Comparative Study" Information 17, no. 1: 9. https://doi.org/10.3390/info17010009
APA StyleUdoidiok, I., Li, F., & Zhang, J. (2026). Evaluating Model Resilience to Data Poisoning Attacks: A Comparative Study. Information, 17(1), 9. https://doi.org/10.3390/info17010009


