paper-with-me

홈 › Papers

Evaluating the Validity of Word-level Adversarial Attacks with Large Language Models

2024-08-15 · Findings of the Association for Computational Linguistics ACL 2024 8 · Huichi Zhou, Zhaoyang Wang, Hongtao Wang, Dongping Chen, Wenhan Mu, Fangyuan Zhang

Deep neural networks exhibit vulnerability to word-level adversarial attacks in natural language processing. Most of these attack methods adopt synonymous substitutions to perturb original samples for crafting adversarial examples while attempting to maintain semantic consistency with the originals. Some of them claim that they could achieve over 90% attack success rate, thereby raising serious safety concerns. However, our investigation reveals that many purportedly successful adversarial examples are actually invalid due to significant changes in semantic meanings compared to their originals. Even when equipped with semantic constraints such as BERTScore, existing attack methods can generate up to 87.9% invalid adversarial examples. Building on this insight, we first curate a 13K dataset for adversarial validity evaluation with the help of GPT-4. Then, an open-source large language model is fine-tuned to offer an interpretable validity score for assessing the semantic consistency between original and adversarial examples. Finally, this validity score can serve as a guide for existing adversarial attack methods to generate valid adversarial examples. Comprehensive experiments demonstrate the effectiveness of our method in evaluating and refining the quality of adversarial examples.

📄 PDF Abstract BibTeX

Code (1)

HuichiZhou/AVLLM

Tasks

Adversarial AttackLanguage ModelingLanguage ModellingLarge Language Modelvalid

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Multi-Head Attention 설명 없음
Adam 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Position-Wise Feed-Forward Layer 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…

Similar Papers 제목 키워드 기반

Adversarial Examples for Evaluating Math Word Problem Solvers

2021-09-13 · Findings (EMNLP) 2021 11 · Vivek Kumar, Rishabh Maheshwary, Vikram Pudi

Standard accuracy metrics have shown that Math Word Problem (MWP) solvers have achieved high performance on benchmark datasets. However, the extent to which existing MWP solvers truly understand language and its relation…

Adversarial RobustnessMathMath Word Problem SolvingSentence

Evaluating Neural Machine Comprehension Model Robustness to Noisy Inputs and Adversarial Attacks

2020-05-01 · Winston Wu, Dustin Arendt, Svitlana Volkova

We evaluate machine comprehension models' robustness to noise and adversarial attacks by performing novel perturbations at the character, word, and sentence level. We experiment with different amounts of perturbations to…

Reading ComprehensionSentence

Evaluating Neural Model Robustness for Machine Comprehension

2021-04-01 · EACL 2021 2 · Winston Wu, Dustin Arendt, Svitlana Volkova

We evaluate neural model robustness to adversarial attacks using different types of linguistic unit perturbations {--} character and word, and propose a new method for strategic sentence-level perturbations. We experimen…

Adversarial AttackmodelReading ComprehensionSentence+1

AtomEval: Validity-Aware Atomic Evaluation of Adversarial Claim Rewriting in Fact Verification

2026-04-09 · Hongyi Cen, Mingxin Wang, Yule Liu, Jingyi Zheng 외 arxiv

Large language models (LLMs) can rewrite refuted claims to evade evidence-based fact verifiers, but conventional attack success rate (ASR) can be inflated when rewrites change, weaken, or correct the false proposition th…

Fact Verification

Defense of Word-level Adversarial Attacks via Random Substitution Encoding

2020-05-01 · Zhao-Yang Wang, Hongtao Wang

The adversarial attacks against deep neural networks on computer vision tasks have spawned many new technologies that help protect models from avoiding false predictions. Recently, word-level adversarial attacks on deep …

General ClassificationSentiment AnalysisSentiment Classificationtext-classification+1