paper-with-me

홈 › Papers

EncT5: Fine-tuning T5 Encoder for Discriminative Tasks

2021-08-17 · ACL ARR August 2021 8 · Anonymous

Encoder-decoder transformer architectures have become popular recently with the advent of T5 models. While they demonstrate impressive performance on benchmarks such as GLUE (Wang et al., 2019), it is not clearly evident if the proposed encoder-decoder architecture is the most efficient for fine-tuning on downstream discriminative tasks. In this work, we study fine-tuning pre-trained encoderdecoder models such as T5. Particularly, we propose EncT5 as a way to efficiently finetune pre-trained encoder-decoder T5 models for classification and regression tasks by using only the encoder layers. Our experimental results show that EncT5 with less than half of the parameters of T5 performs similarly to T5 models on GLUE benchmark. We believe our proposed approach can be easily applied to any pre-trained encoder-decoder model.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Decoder

Methods 이 논문이 사용한 방법론

Gated Linear Unit A Gated Linear Unit, or GLU computes: $$ \mathrm{GLU}(a, b) = a \otimes \sigma(b) $$ It is used in natural language processing architectures, for example the Gated CNN,…
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Inverse Square Root Schedule Inverse Square Root is a learning rate schedule 1 / $\sqrt{\max\left(n, k\right)}$ where $n$ is the current training iteration and $k$ is the number of warm-up steps. This…
Adafactor Adafactor is a stochastic optimization method based on Adam that reduces memory usage while retaining the empirical benefits of…

Similar Papers 제목 키워드 기반

EncT5: A Framework for Fine-tuning T5 as Non-autoregressive Models

2021-10-16 · Frederick Liu, Terry Huang, Shihang Lyu, Siamak Shakeri 외

Pre-trained encoder-decoder transformer architectures have become increasingly popular recently with the advent of T5 models. T5 has also become more favorable over other architectures like BERT due to the amount of data…

DecoderLanguage ModellingMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION+2

A New Probabilistic V-Net Model with Hierarchical Spatial Feature Transform for Efficient Abdominal Multi-Organ Segmentation

2022-08-02 · Minfeng Xu, Heng Guo, Jianfeng Zhang, Ke Yan 외

Accurate and robust abdominal multi-organ segmentation from CT imaging of different modalities is a challenging task due to complex inter- and intra-organ shape and appearance variations among abdominal organs. In this p…

DecoderOrgan SegmentationSegmentation

AbdomenCT-1K: Is Abdominal Organ Segmentation A Solved Problem?

2020-10-28 · Jun Ma, Yao Zhang, Song Gu, Cheng Zhu 외

With the unprecedented developments in deep learning, automatic segmentation of main abdominal organs seems to be a solved problem as state-of-the-art (SOTA) methods have achieved comparable results with inter-rater vari…

Continual LearningOrgan SegmentationPancreas SegmentationSegmentation

Efficient and Versatile Robust Fine-Tuning of Zero-shot Models

2024-08-11 · Sungyeon Kim, Boseung Jeong, Donghyun Kim, Suha Kwak

Large-scale image-text pre-trained models enable zero-shot classification and provide consistent accuracy across various data distributions. Nonetheless, optimizing these models in downstream tasks typically requires fin…

Cross-Modal Retrievalzero-shot-classificationZero-Shot Learning

Pre-Calc: Learning to Use the Calculator Improves Numeracy in Language Models

2024-04-22 · Vishruth Veerendranath, Vishwa Shah, Kshitish Ghate

Quantitative and numerical comprehension in language is an important task in many fields like education and finance, but still remains a challenging task for language models. While tool and calculator usage has shown to …

DecoderMathematical Reasoning