paper-with-me

홈 › Papers

Prompting ELECTRA: Few-Shot Learning with Discriminative Pre-Trained Models

2022-05-30 · Mengzhou Xia, Mikel Artetxe, Jingfei Du, Danqi Chen, Ves Stoyanov

Pre-trained masked language models successfully perform few-shot learning by formulating downstream tasks as text infilling. However, as a strong alternative in full-shot settings, discriminative pre-trained models like ELECTRA do not fit into the paradigm. In this work, we adapt prompt-based few-shot learning to ELECTRA and show that it outperforms masked language models in a wide range of tasks. ELECTRA is pre-trained to distinguish if a token is generated or original. We naturally extend that to prompt-based few-shot learning by training to score the originality of the target options without introducing new parameters. Our method can be easily adapted to tasks involving multi-token predictions without extra computation overhead. Analysis shows that ELECTRA learns distributions that align better with downstream tasks.

📄 PDF Abstract BibTeX arXiv:2205.15223

Code (1)

facebookresearch/electra-fewshot-learning 공식 구현 pytorch

Tasks

Few-Shot LearningText Infilling

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…

Similar Papers 제목 키워드 기반

ELECTRA is a Zero-Shot Learner, Too

2022-07-17 · Shiwen Ni, Hung-Yu Kao

Recently, for few-shot or even zero-shot learning, the new paradigm "pre-train, prompt, and predict" has achieved remarkable achievements compared with the "pre-train, fine-tune" paradigm. After the success of prompt-bas…

Language ModelingLanguage ModellingPrompt LearningSST-2+1

Discriminative Language Model as Semantic Consistency Scorer for Prompt-based Few-Shot Text Classification

2022-10-23 · Zhipeng Xie, Yahe Li

This paper proposes a novel prompt-based finetuning method (called DLM-SCS) for few-shot text classification by utilizing the discriminative language model ELECTRA that is pretrained to distinguish whether a token is ori…

Few-Shot Text ClassificationLanguage ModelingLanguage Modellingtext-classification+1

On the effectiveness of small, discriminatively pre-trained language representation models for biomedical text mining

2020-11-01 · EMNLP (sdp) 2020 11 · Ibrahim Burak Ozyurt

Neural language representation models such as BERT have recently shown state of the art performance in downstream NLP tasks and bio-medical domain adaptation of BERT (Bio-BERT) has shown same behavior on biomedical text …

Domain AdaptationGPUnamed-entity-recognitionNamed Entity Recognition+3

Vita-CLIP: Video and text adaptive CLIP via Multimodal Prompting

2023-04-06 · CVPR 2023 1 · Syed Talal Wasim, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan 외

Adopting contrastive image-text pretrained models like CLIP towards video classification has gained attention due to its cost-effectiveness and competitive performance. However, recent works in this area face a trade-off…

Action RecognitionPrompt LearningVideo ClassificationZero-Shot Action Recognition+1

Prompt Tuning for Discriminative Pre-trained Language Models

2022-05-23 · Findings (ACL) 2022 5 · Yuan YAO, Bowen Dong, Ao Zhang, Zhengyan Zhang 외

Recent works have shown promising results of prompt tuning in stimulating pre-trained language models (PLMs) for natural language processing (NLP) tasks. However, to the best of our knowledge, existing works focus on pro…

Language ModelingLanguage ModellingQuestion Answeringtext-classification+1