paper-with-me

홈 › Papers

Pre-trained Token-replaced Detection Model as Few-shot Learner

2022-03-07 · COLING 2022 10 · Zicheng Li, Shoushan Li, Guodong Zhou

Pre-trained masked language models have demonstrated remarkable ability as few-shot learners. In this paper, as an alternative, we propose a novel approach to few-shot learning with pre-trained token-replaced detection models like ELECTRA. In this approach, we reformulate a classification or a regression task as a token-replaced detection problem. Specifically, we first define a template and label description words for each task and put them into the input to form a natural language prompt. Then, we employ the pre-trained token-replaced detection model to predict which label description word is the most original (i.e., least replaced) among all label description words in the prompt. A systematic evaluation on 16 datasets demonstrates that our approach outperforms few-shot learners with pre-trained masked language models in both one-sentence and two-sentence learning tasks.

📄 PDF Abstract BibTeX arXiv:2203.03235

Code (1)

cjfarmer/trd_fsl 공식 구현 pytorch

Tasks

Few-Shot LearningSentence

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Adam 설명 없음
WordPiece 설명 없음
Weight Decay 설명 없음

Similar Papers 제목 키워드 기반

ELECTRA is a Zero-Shot Learner, Too

2022-07-17 · Shiwen Ni, Hung-Yu Kao

Recently, for few-shot or even zero-shot learning, the new paradigm "pre-train, prompt, and predict" has achieved remarkable achievements compared with the "pre-train, fine-tune" paradigm. After the success of prompt-bas…

Language ModelingLanguage ModellingPrompt LearningSST-2+1

Assessing Phrase Break of ESL speech with Pre-trained Language Models

2022-10-28 · Zhiyi Wang, Shaoguang Mao, Wenshan Wu, Yan Xia

This work introduces an approach to assessing phrase break in ESL learners' speech with pre-trained language models (PLMs). Different with traditional methods, this proposal converts speech to token sequences, and then l…

text-classificationText Classification

Assessing Phrase Break of ESL Speech with Pre-trained Language Models and Large Language Models

2023-06-08 · Zhiyi Wang, Shaoguang Mao, Wenshan Wu, Yan Xia 외

This work introduces approaches to assessing phrase breaks in ESL learners' speech using pre-trained language models (PLMs) and large language models (LLMs). There are two tasks: overall assessment of phrase break for a …

text-classificationText Classification

X-PuDu at SemEval-2022 Task 7: A Replaced Token Detection Task Pre-trained Model with Pattern-aware Ensembling for Identifying Plausible Clarifications

2022-11-27 · SemEval (NAACL) 2022 7 · Junyuan Shang, Shuohuan Wang, Yu Sun, Yanjun Yu 외

This paper describes our winning system on SemEval 2022 Task 7: Identifying Plausible Clarifications of Implicit and Underspecified Phrases in Instructional Texts. A replaced token detection pre-trained model is utilized…

Multi-class Classification

GanLM: Encoder-Decoder Pre-training with an Auxiliary Discriminator

2022-12-20 · Jian Yang, Shuming Ma, Li Dong, Shaohan Huang 외

Pre-trained models have achieved remarkable success in natural language processing (NLP). However, existing pre-training methods underutilize the benefits of language understanding for generation. Inspired by the idea of…

DecoderDenoisingSentenceText Generation