Using Random Perturbations to Mitigate Adversarial Attacks on Sentiment Analysis Models
Attacks on deep learning models are often difficult to identify and therefore are difficult to protect against. This problem is exacerbated by the use of public datasets that typically are not manually inspected before use. In this paper, we offer a solution to this vulnerability by using, during testing, random perturbations such as spelling correction if necessary, substitution by random synonym, or simply dropping the word. These perturbations are applied to random words in random sentences to defend NLP models against adversarial attacks. Our Random Perturbations Defense and Increased Randomness Defense methods are successful in returning attacked models to similar accuracy of models before attacks. The original accuracy of the model used in this work is 80% for sentiment classification. After undergoing attacks, the accuracy drops to accuracy between 0% and 44%. After applying our defense methods, the accuracy of the model is returned to the original accuracy within statistical significance.
Code (0)
등록된 구현이 없습니다.
Tasks
Sentiment AnalysisSentiment ClassificationSpelling CorrectionSimilar Papers 제목 키워드 기반
Learning to Discriminate Perturbations for Blocking Adversarial Attacks in Text Classification
Adversarial attacks against machine learning models have threatened various real-world applications such as spam filtering and sentiment analysis. In this paper, we propose a novel framework, learning to DIScriminate Per…
BlockingGeneral ClassificationSentiment Analysistext-classification+1Adversarial Evasion Attack Efficiency against Large Language Models
Large Language Models (LLMs) are valuable for text classification, but their vulnerabilities must not be disregarded. They lack robustness against adversarial examples, so it is pertinent to understand the impacts of dif…
Adversarial DefenseClassificationSentiment AnalysisSentiment Classification+2A Mask-Based Adversarial Defense Scheme
Adversarial attacks hamper the functionality and accuracy of Deep Neural Networks (DNNs) by meddling with subtle perturbations to their inputs.In this work, we propose a new Mask-based Adversarial Defense scheme (MAD) fo…
Adversarial AttackAdversarial DefenseDenoisingEnd-to-End Adversarial White Box Attacks on Music Instrument Classification
Small adversarial perturbations of input data are able to drastically change performance of machine learning systems, thereby challenging the validity of such systems. We present the very first end-to-end adversarial att…
BIG-bench Machine LearningGeneral ClassificationHow adversarial attacks can disrupt seemingly stable accurate classifiers
Adversarial attacks dramatically change the output of an otherwise accurate learning system using a seemingly inconsequential modification to a piece of input data. Paradoxically, empirical evidence indicates that even s…
image-classificationImage Classification