paper-with-me

홈 › Papers

Adversarial Self-Attention for Language Understanding

2022-06-25 · Hongqiu Wu, Ruixue Ding, Hai Zhao, Pengjun Xie, Fei Huang, Min Zhang

Deep neural models (e.g. Transformer) naturally learn spurious features, which create a ``shortcut'' between the labels and inputs, thus impairing the generalization and robustness. This paper advances the self-attention mechanism to its robust variant for Transformer-based pre-trained language models (e.g. BERT). We propose \textit{Adversarial Self-Attention} mechanism (ASA), which adversarially biases the attentions to effectively suppress the model reliance on features (e.g. specific keywords) and encourage its exploration of broader semantics. We conduct a comprehensive evaluation across a wide range of tasks for both pre-training and fine-tuning stages. For pre-training, ASA unfolds remarkable performance gains compared to naive training for longer steps. For fine-tuning, ASA-empowered models outweigh naive models by a large margin considering both generalization and robustness.

📄 PDF Abstract BibTeX arXiv:2206.12608

Code (1)

gingasan/adversarialsa 공식 구현 pytorch

Tasks

Machine Reading ComprehensionNamed Entity Recognition (NER)Natural Language InferenceParaphrase IdentificationSemantic SimilaritySemantic Textual SimilaritySentiment Analysis

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Adam 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…

Similar Papers 제목 키워드 기반

Alignment Attention by Matching Key and Query Distributions

2021-10-25 · NeurIPS 2021 12 · Shujian Zhang, Xinjie Fan, Huangjie Zheng, Korawat Tanwisuth 외

The neural attention mechanism has been incorporated into deep neural networks to achieve state-of-the-art performance in various domains. Most such models use multi-head self-attention which is appealing for the ability…

Graph AttentionQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Fooling Vision and Language Models Despite Localization and Attention Mechanism

2017-09-25 · CVPR 2018 6 · Xiaojun Xu, Xinyun Chen, Chang Liu, Anna Rohrbach 외

Adversarial attacks are known to succeed on classifiers, but it has been an open question whether more complex vision systems are vulnerable. In this paper, we study adversarial examples for vision and language models, w…

Dense CaptioningNatural Language UnderstandingOpen-Ended Question AnsweringQuestion Answering+2

Enhancing Essay Scoring with Adversarial Weights Perturbation and Metric-specific AttentionPooling

2024-01-06 · Jiaxin Huang, Xinyu Zhao, Chang Che, Qunwei Lin 외

The objective of this study is to improve automated feedback tools designed for English Language Learners (ELLs) through the utilization of data science techniques encompassing machine learning, natural language processi…

Automated Essay ScoringLanguage ModellingNatural Language UnderstandingSelf-Supervised Learning

SALSA-TEXT : self attentive latent space based adversarial text generation

2018-09-28 · Jules Gagnon-Marchand, Hamed Sadeghi, Md. Akmal Haidar, Mehdi Rezagholizadeh

Inspired by the success of self attention mechanism and Transformer architecture in sequence transduction and image generation applications, we propose novel self attention-based architectures to improve the performance …

Adversarial TextImage GenerationSentenceSentence Compression+1

Towards Efficient Adversarial Training on Vision Transformers

2022-07-21 · Boxi Wu, Jindong Gu, Zhifeng Li, Deng Cai 외

Vision Transformer (ViT), as a powerful alternative to Convolutional Neural Network (CNN), has received much attention. Recent work showed that ViTs are also vulnerable to adversarial examples like CNNs. To build robust …