Next Article in Journal
Synchronized Multi-Augmentation with Multi-Backbone Ensembling for Enhancing Deep Learning Performance
Previous Article in Journal
Investigating Stress During a Virtual Reality Game Through Fractal and Multifractal Analysis of Heart Rate Variability
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Explainable Artificial Intelligence with Integrated Gradients for the Detection of Adversarial Attacks on Text Classifiers

by
Harsha Moraliyage
1,
Geemini Kulawardana
1,
Daswin De Silva
1,*,
Zafar Issadeen
1,
Milos Manic
2 and
Seiichiro Katsura
3
1
Centre for Data Analytics and Cognition, La Trobe University, Melbourne, VIC 3083, Australia
2
Department of Computer Science, Virginia Commonwealth University, Richmond, VA 23284-2520, USA
3
Department of System Design Engineering, Keio University, Tokyo 108-8345, Japan
*
Author to whom correspondence should be addressed.
Appl. Syst. Innov. 2025, 8(1), 17; https://doi.org/10.3390/asi8010017
Submission received: 20 September 2024 / Revised: 9 January 2025 / Accepted: 10 January 2025 / Published: 21 January 2025
(This article belongs to the Section Artificial Intelligence)

Abstract

Text classifiers are Artificial Intelligence (AI) models used to classify new documents or text vectors into predefined classes. They are typically built using supervised learning algorithms and labelled datasets. Text classifiers produce a predefined class as an output, which also makes them susceptible to adversarial attacks. Text classifiers with high accuracy that are trained using complex deep learning algorithms are equally susceptible to adversarial examples, due to subtle differences that are indiscernible to human experts. Recent work in this space is mostly focused on improving adversarial robustness and adversarial example detection, instead of detecting adversarial attacks. In this paper, we propose a novel approach, explainable AI with integrated gradients (IGs) for the detection of adversarial attacks on text classifiers. This approach uses IGs to unpack model behavior and identify terms that positively and negatively influence the target prediction. Instead of random substitution of words in the input, we select the top p% words with the greatest positive and negative influence as substitute candidates using attribution scores obtained from IGs to generate k samples of transformed inputs by replacing them with synonyms. This approach does not require changes to the model architecture or the training algorithm. The approach was empirically evaluated on three benchmark datasets, IMDB, SST-2, and AG News. Our approach outperforms baseline models on word substitution rate, detection accuracy, and F1 scores while maintaining equivalent detection performance against adversarial attacks.
Keywords: adversarial attacks; AI cybersecurity; integrated gradients; text classification; explainable AI adversarial attacks; AI cybersecurity; integrated gradients; text classification; explainable AI

Share and Cite

MDPI and ACS Style

Moraliyage, H.; Kulawardana, G.; De Silva, D.; Issadeen, Z.; Manic, M.; Katsura, S. Explainable Artificial Intelligence with Integrated Gradients for the Detection of Adversarial Attacks on Text Classifiers. Appl. Syst. Innov. 2025, 8, 17. https://doi.org/10.3390/asi8010017

AMA Style

Moraliyage H, Kulawardana G, De Silva D, Issadeen Z, Manic M, Katsura S. Explainable Artificial Intelligence with Integrated Gradients for the Detection of Adversarial Attacks on Text Classifiers. Applied System Innovation. 2025; 8(1):17. https://doi.org/10.3390/asi8010017

Chicago/Turabian Style

Moraliyage, Harsha, Geemini Kulawardana, Daswin De Silva, Zafar Issadeen, Milos Manic, and Seiichiro Katsura. 2025. "Explainable Artificial Intelligence with Integrated Gradients for the Detection of Adversarial Attacks on Text Classifiers" Applied System Innovation 8, no. 1: 17. https://doi.org/10.3390/asi8010017

APA Style

Moraliyage, H., Kulawardana, G., De Silva, D., Issadeen, Z., Manic, M., & Katsura, S. (2025). Explainable Artificial Intelligence with Integrated Gradients for the Detection of Adversarial Attacks on Text Classifiers. Applied System Innovation, 8(1), 17. https://doi.org/10.3390/asi8010017

Article Metrics

Back to TopTop