paper-with-me

홈 › Papers

ELECTRA is a Zero-Shot Learner, Too

2022-07-17 · Shiwen Ni, Hung-Yu Kao

Recently, for few-shot or even zero-shot learning, the new paradigm "pre-train, prompt, and predict" has achieved remarkable achievements compared with the "pre-train, fine-tune" paradigm. After the success of prompt-based GPT-3, a series of masked language model (MLM)-based (e.g., BERT, RoBERTa) prompt learning methods became popular and widely used. However, another efficient pre-trained discriminative model, ELECTRA, has probably been neglected. In this paper, we attempt to accomplish several NLP tasks in the zero-shot scenario using a novel our proposed replaced token detection (RTD)-based prompt learning method. Experimental results show that ELECTRA model based on RTD-prompt learning achieves surprisingly state-of-the-art zero-shot performance. Numerically, compared to MLM-RoBERTa-large and MLM-BERT-large, our RTD-ELECTRA-large has an average of about 8.4% and 13.7% improvement on all 15 tasks. Especially on the SST-2 task, our RTD-ELECTRA-large achieves an astonishing 90.1% accuracy without any training data. Overall, compared to the pre-trained masked language models, the pre-trained replaced token detection model performs better in zero-shot learning. The source code is available at: https://github.com/nishiwen1214/RTD-ELECTRA.

📄 PDF Abstract BibTeX arXiv:2207.08141

Code (1)

nishiwen1214/rtd-electra 공식 구현 tf

Tasks

Language ModelingLanguage ModellingPrompt LearningSST-2Zero-Shot Learning

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
{Dispute@FaQ-s}How to file a dispute with Expedia? How to file a dispute with Expedia? To file a complaint against Expedia, first try contacting their customer service directly. You can reach them by phone at…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Weight Decay 설명 없음

Similar Papers 제목 키워드 기반

ELECTRA and GPT-4o: Cost-Effective Partners for Sentiment Analysis

2024-12-29 · James P. Beno

Bidirectional transformers excel at sentiment analysis, and Large Language Models (LLM) are effective zero-shot learners. Might they perform better as a team? This paper explores collaborative approaches between ELECTRA …

Sentiment AnalysisSentiment Classification

Pre-trained Token-replaced Detection Model as Few-shot Learner

2022-03-07 · COLING 2022 10 · Zicheng Li, Shoushan Li, Guodong Zhou

Pre-trained masked language models have demonstrated remarkable ability as few-shot learners. In this paper, as an alternative, we propose a novel approach to few-shot learning with pre-trained token-replaced detection m…

Few-Shot LearningSentence

Prompting ELECTRA: Few-Shot Learning with Discriminative Pre-Trained Models

2022-05-30 · Mengzhou Xia, Mikel Artetxe, Jingfei Du, Danqi Chen 외

Pre-trained masked language models successfully perform few-shot learning by formulating downstream tasks as text infilling. However, as a strong alternative in full-shot settings, discriminative pre-trained models like …

Few-Shot LearningText Infilling

When Informal Text Breaks NLI: Tokenization Failure, Distribution Shift, and Targeted Mitigations

2026-04-18 · Avinash Goutham Aluguvelly arxiv

We study how informal surface forms degrade NLI accuracy in ELECTRA-small (14M) and RoBERTa-large (355M) across four transforms applied to SNLI and MultiNLI: slang substitution, emoji replacement, Gen-Z filler tokens, an…

Projected Subnetworks Scale Adaptation

2023-01-27 · Siddhartha Datta, Nigel Shadbolt

Large models support great zero-shot and few-shot capabilities. However, updating these models on new tasks can break performance on previous seen tasks and their zero/few-shot unseen tasks. Our work explores how to upda…