paper-with-me

홈 › Papers

Domain Adaptation from Scratch

2022-09-02 · Eyal Ben-David, Yftah Ziser, Roi Reichart

Natural language processing (NLP) algorithms are rapidly improving but often struggle when applied to out-of-distribution examples. A prominent approach to mitigate the domain gap is domain adaptation, where a model trained on a source domain is adapted to a new target domain. We present a new learning setup, ``domain adaptation from scratch'', which we believe to be crucial for extending the reach of NLP to sensitive domains in a privacy-preserving manner. In this setup, we aim to efficiently annotate data from a set of source domains such that the trained model performs well on a sensitive target domain from which data is unavailable for annotation. Our study compares several approaches for this challenging setup, ranging from data selection and domain adaptation algorithms to active learning paradigms, on two NLP tasks: sentiment analysis and Named Entity Recognition. Our results suggest that using the abovementioned approaches eases the domain gap, and combining them further improves the results.

📄 PDF Abstract BibTeX arXiv:2209.00830

Code (1)

eyalbd2/scratchda 공식 구현

Tasks

Active LearningDomain Adaptationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Privacy PreservingSentiment Analysis

Similar Papers 제목 키워드 기반

From Scratch to Fine-Tuned: A Comparative Study of Transformer Training Strategies for Legal Machine Translation

2025-12-21 · Amit Barman, Atanu Mandal, Sudip Kumar Naskar arxiv

In multilingual nations like India, access to legal information is often hindered by language barriers, as much of the legal and judicial documentation remains in English. Legal Machine Translation (L-MT) offers a scalab…

Machine TranslationDomain Adaptation

Overcoming Concept Shift in Domain-Aware Settings through Consolidated Internal Distributions

2020-07-01 · Mohammad Rostami, Aram Galstyan

We develop an algorithm to improve the performance of a pre-trained model under concept shift without retraining the model from scratch when only unannotated samples of initial concepts are accessible. We model this prob…

Domain AdaptationTransfer LearningUnsupervised Domain Adaptation

Legal Domain Adaptation of Modern BERT Models

2026-06-26 · Dominik Stammbach, Peter Henderson arxiv

We investigate domain adaptation of modern BERT models in the legal domain. We further pre-train ModernBERT on all US court opinions using the masked language modeling objective. Although ModernBERT has been trained on r…

Domain Adaptation

The Word and the Way: Strategies for Domain-Specific BERT Pre-Training in German Medical NLP

2026-06-02 · Henry He, Johann Frei, Raphael Schmitt arxiv

Digital healthcare generates vast amounts of clinical text that can support AI-assisted applications, yet German biomedical language models remain limited by older architectures or restricted training data. We present Ch…

Medical Named Entity RecognitionText ClassificationDomain Adaptation

NeuroADDA: Active Discriminative Domain Adaptation in Connectomic

2025-03-08 · Shashata Sawmya, Thomas L. Athey, Gwyneth Liu, Nir Shavit

Training segmentation models from scratch has been the standard approach for new electron microscopy connectomics datasets. However, leveraging pretrained models from existing datasets could improve efficiency and perfor…

Active LearningDomain AdaptationTransfer Learning