paper-with-me

홈 › Papers

How Useful is Continued Pre-Training for Generative Unsupervised Domain Adaptation?

2024-01-31 · Rheeya Uppaal, Yixuan Li, Junjie Hu

Recent breakthroughs in scale have enabled the emergence of powerful generative language models, and the ability to fine-tune these models on various tasks by casting them into prompts or instructions. In this landscape, the problem of Unsupervised Domain Adaptation (UDA), or the problem of leveraging knowledge from a labeled source domain to an unlabeled target domain, has been left behind, with recent UDA methods still addressing discriminative classification. In particular, two popular UDA approaches, involving Continued Pre-Training (CPT) and learning domain invariant representations, have been under-explored in the generative setting, signaling a gap. In this work, we evaluate the utility of CPT for generative UDA. We first perform an empirical evaluation to measure the trade-offs between CPT and strong methods promoting domain invariance. We further evaluate how well the benefits of CPT extend to different architectures, tuning methods and data regimes. We then motivate the use of CPT by studying to what degree it benefits classification performance on the target domain. Finally, we attempt to understand the mechanism behind which CPT improves classification performance on the unlabeled target domain. Our findings suggest that a implicitly learns the downstream task while predicting masked words informative to that task. Our work connects the body of UDA research with that of instruction tuning, enabling an initial step towards a wider applicability of modern language models.

📄 PDF Abstract BibTeX arXiv:2401.17514

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationDomain Adaptationdomain classificationLanguage ModellingMasked Language ModelingSelf-Supervised LearningUnsupervised Domain Adaptation

Similar Papers 제목 키워드 기반

TapWeight: Reweighting Pretraining Objectives for Task-Adaptive Pretraining

2024-10-13 · Ruiyi Zhang, Sai Ashish Somayajula, Pengtao Xie

Large-scale general domain pretraining followed by downstream-specific finetuning has become a predominant paradigm in machine learning. However, discrepancies between the pretraining and target domains can still lead to…

Molecular Property PredictionNatural Language UnderstandingProperty Prediction

Customized Generative AI Agent for Transportation Engineering Practice: A Development and Continued Pre-training Guideline

2026-06-27 · Dianwei Chen, Yuan-Zheng Lei, Zifan Zhang, Yuchen Liu 외 arxiv

Recent advancements in generative artificial intelligence (AI) and large language models (LLMs) have shown significant promise in automating complex reasoning, summarization, and question-answering tasks. However, the ef…

Stable Distillation: Regularizing Continued Pre-training for Low-Resource Automatic Speech Recognition

2023-12-20 · Ashish Seth, Sreyan Ghosh, S. Umesh, Dinesh Manocha

Continued self-supervised (SSL) pre-training for adapting existing SSL models to the target domain has shown to be extremely effective for low-resource Automatic Speech Recognition (ASR). This paper proposes Stable Disti…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Crossing-Domain Generative Adversarial Networks for Unsupervised Multi-Domain Image-to-Image Translation

2020-08-27 · Xuewen Yang, Dongliang Xie, Xin Wang

State-of-the-art techniques in Generative Adversarial Networks (GANs) have shown remarkable success in image-to-image translation from peer domain X to domain Y using paired image data. However, obtaining abundant paired…

Image-to-Image TranslationTranslationUnsupervised Image-To-Image Translation

IGOT: Information Gain Optimized Tokenizer on Domain Adaptive Pretraining

2024-05-16 · Dawei Feng, Yihai Zhang, Zhixuan Xu

Pretrained Large Language Models (LLM) such as ChatGPT, Claude, etc. have demonstrated strong capabilities in various fields of natural language generation. However, there are still many problems when using LLM in specia…

Domain AdaptationGPUText Generation