paper-with-me

홈 › Papers

Back-Translated Task Adaptive Pretraining: Improving Accuracy and Robustness on Text Classification

2021-07-22 · Junghoon Lee, Jounghee Kim, Pilsung Kang

Language models (LMs) pretrained on a large text corpus and fine-tuned on a downstream text corpus and fine-tuned on a downstream task becomes a de facto training strategy for several natural language processing (NLP) tasks. Recently, an adaptive pretraining method retraining the pretrained language model with task-relevant data has shown significant performance improvements. However, current adaptive pretraining methods suffer from underfitting on the task distribution owing to a relatively small amount of data to re-pretrain the LM. To completely use the concept of adaptive pretraining, we propose a back-translated task-adaptive pretraining (BT-TAPT) method that increases the amount of task-specific data for LM re-pretraining by augmenting the task data using back-translation to generalize the LM to the target task domain. The experimental results show that the proposed BT-TAPT yields improved classification accuracy on both low- and high-resource data and better robustness to noise than the conventional adaptive pretraining method.

📄 PDF Abstract BibTeX arXiv:2107.10474

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modellingtext-classificationText ClassificationTranslation

Similar Papers 제목 키워드 기반

Multilingual Multimodal Learning with Machine Translated Text

2022-10-24 · Chen Qiu, Dan Oneata, Emanuele Bugliarello, Stella Frank 외

Most vision-and-language pretraining research focuses on English tasks. However, the creation of multilingual multimodal evaluation datasets (e.g. Multi30K, xGQA, XVNLI, and MaRVL) poses a new challenge in finding high-q…

Zero-Shot Cross-Lingual Image-to-Text RetrievalZero-Shot Cross-Lingual Text-to-Image RetrievalZero-Shot Cross-Lingual Visual Natural Language InferenceZero-Shot Cross-Lingual Visual Question Answering+1

Detecting Machine-Translated Text using Back Translation

2019-10-15 · WS 2019 10 · Hoang-Quoc Nguyen-Son, Tran Phuong Thao, Seira Hidano, Shinsaku Kiyomoto

Machine-translated text plays a crucial role in the communication of people using different languages. However, adversaries can use such text for malicious purposes such as plagiarism and fake review. The existing method…

Translation

Adobe AMPS’s Submission for Very Low Resource Supervised Translation Task at WMT20

2020-11-01 · WMT (EMNLP) 2020 11 · Keshaw Singh

In this paper, we describe our systems submitted to the very low resource supervised translation task at WMT20. We participate in both translation directions for Upper Sorbian-German language pair. Our primary submission…

Machine TranslationTranslation

TapWeight: Reweighting Pretraining Objectives for Task-Adaptive Pretraining

2024-10-13 · Ruiyi Zhang, Sai Ashish Somayajula, Pengtao Xie

Large-scale general domain pretraining followed by downstream-specific finetuning has become a predominant paradigm in machine learning. However, discrepancies between the pretraining and target domains can still lead to…

Molecular Property PredictionNatural Language UnderstandingProperty Prediction

Learning What to Predict: Downstream-Guided Task Design for Continued Pretraining

2026-01-29 · Shuqi Ke, Giulia Fanti arxiv

Continued pretraining is optimized with fixed self-supervised tasks but selected by downstream performance, creating a coarse feedback loop in which practitioners evaluate checkpoints, change data mixtures or objectives,…

Semantic SegmentationDepth Estimation