paper-with-me

홈 › Papers

GradMask: Gradient-Guided Token Masking for Textual Adversarial Example Detection

2021-11-16 · ACL ARR September 2021 9 · Anonymous

We present a simple model-agnostic textual adversarial example detection scheme called GradMask. It uses gradient signals to detect adversarially perturbed tokens in an input sequence and occludes such tokens by a masking process. GradMask provides several advantages over existing methods including improved detection performance and a weak interpretation of its decision. Extensive evaluations on widely adopted natural language processing benchmark datasets demonstrate the efficiency and effectiveness of GradMask

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GradMask: Effective Fine-tuning on Large-scale Pretrained Language Models via Gradient Masking

2021-05-24 · Anonymous

Pretrained language models have dominated a variety of NLP tasks. However, fine-tuning large pretrained models on downstream tasks tend to achieve degenerated and unstable results, especially when there are only a limite…

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval

2025-09-11 · Tianlu Zheng, Yifan Zhang, Xiang An, Ziyong Feng 외 arxiv

Although Contrastive Language-Image Pre-training (CLIP) exhibits strong performance across diverse vision tasks, its application to person representation learning faces two critical challenges: (i) the scarcity of large-…

Representation LearningContrastive LearningPerson Retrieval

GradMask: Reduce Overfitting by Regularizing Saliency

2019-04-16 · Becks Simpson, Francis Dutil, Yoshua Bengio, Joseph Paul Cohen

With too few samples or too many model parameters, overfitting can inhibit the ability to generalise predictions to new data. Within medical imaging, this can occur when features are incorrectly assigned importance such …

Lesion Segmentation

Attribution-Guided Masking for Robust Cross-Domain Sentiment Classification

2026-05-04 · Shubham Harkare, Arvind Yogesh Suresh Babu, Yash Kulkarni arxiv

While pre-trained Transformer models achieve high accuracy on in-domain sentiment classification, they frequently experience severe performance degradation when transferring to out-of-domain data. We hypothesize that thi…

Localize and Neutralize: Gradient-guided Token Suppression against Visual Prompt Injection Attack

2026-05-24 · Dongpeng Zhang, Ke Ma, Yangbangyan Jiang, Gaozheng Pei 외 arxiv

Adversarial images pose a severe security threat to multimodal large language models through prompt injection. Existing defenses largely lack a principled understanding of the underlying mechanisms and struggle to balanc…

Adversarial Attack