Evaluating the Validity of Word-level Adversarial Attacks with Large Language Models
Deep neural networks exhibit vulnerability to word-level adversarial attacks in natural language processing. Most of these attack methods adopt synonymous substitutions to perturb original samples for crafting adversarial examples while attempting to maintain semantic consistency with the originals. Some of them claim that they could achieve over 90% attack success rate, thereby raising serious safety concerns. However, our investigation reveals that many purportedly successful adversarial examples are actually invalid due to significant changes in semantic meanings compared to their originals. Even when equipped with semantic constraints such as BERTScore, existing attack methods can generate up to 87.9% invalid adversarial examples. Building on this insight, we first curate a 13K dataset for adversarial validity evaluation with the help of GPT-4. Then, an open-source large language model is fine-tuned to offer an interpretable validity score for assessing the semantic consistency between original and adversarial examples. Finally, this validity score can serve as a guide for existing adversarial attack methods to generate valid adversarial examples. Comprehensive experiments demonstrate the effectiveness of our method in evaluating and refining the quality of adversarial examples.
Code (1)
Tasks
Adversarial AttackLanguage ModelingLanguage ModellingLarge Language ModelvalidMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Adversarial Examples for Evaluating Math Word Problem Solvers
Standard accuracy metrics have shown that Math Word Problem (MWP) solvers have achieved high performance on benchmark datasets. However, the extent to which existing MWP solvers truly understand language and its relation…
Adversarial RobustnessMathMath Word Problem SolvingSentenceEvaluating Neural Machine Comprehension Model Robustness to Noisy Inputs and Adversarial Attacks
We evaluate machine comprehension models' robustness to noise and adversarial attacks by performing novel perturbations at the character, word, and sentence level. We experiment with different amounts of perturbations to…
Reading ComprehensionSentenceEvaluating Neural Model Robustness for Machine Comprehension
We evaluate neural model robustness to adversarial attacks using different types of linguistic unit perturbations {--} character and word, and propose a new method for strategic sentence-level perturbations. We experimen…
Adversarial AttackmodelReading ComprehensionSentence+1AtomEval: Validity-Aware Atomic Evaluation of Adversarial Claim Rewriting in Fact Verification
Large language models (LLMs) can rewrite refuted claims to evade evidence-based fact verifiers, but conventional attack success rate (ASR) can be inflated when rewrites change, weaken, or correct the false proposition th…
Fact VerificationDefense of Word-level Adversarial Attacks via Random Substitution Encoding
The adversarial attacks against deep neural networks on computer vision tasks have spawned many new technologies that help protect models from avoiding false predictions. Recently, word-level adversarial attacks on deep …
General ClassificationSentiment AnalysisSentiment Classificationtext-classification+1