paper-with-me

홈 › Papers

LegalRelectra: Mixed-domain Language Modeling for Long-range Legal Text Comprehension

2022-12-16 · Wenyue Hua, Yuchen Zhang, Zhe Chen, Josie Li, Melanie Weber

The application of Natural Language Processing (NLP) to specialized domains, such as the law, has recently received a surge of interest. As many legal services rely on processing and analyzing large collections of documents, automating such tasks with NLP tools emerges as a key challenge. Many popular language models, such as BERT or RoBERTa, are general-purpose models, which have limitations on processing specialized legal terminology and syntax. In addition, legal documents may contain specialized vocabulary from other domains, such as medical terminology in personal injury text. Here, we propose LegalRelectra, a legal-domain language model that is trained on mixed-domain legal and medical corpora. We show that our model improves over general-domain and single-domain medical and legal language models when processing mixed-domain (personal injury) text. Our training architecture implements the Electra framework, but utilizes Reformer instead of BERT for its generator and discriminator. We show that this improves the model's performance on processing long passages and results in better long-range text comprehension.

📄 PDF Abstract BibTeX arXiv:2212.08204

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingReading Comprehension

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Reversible Residual Block 설명 없음
Adafactor Adafactor is a stochastic optimization method based on Adam that reduces memory usage while retaining the empirical benefits of…
LSH Attention 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

UDALM: Unsupervised Domain Adaptation through Language Modeling

2021-04-14 · NAACL 2021 4 · Constantinos Karouzos, Georgios Paraskevopoulos, Alexandros Potamianos

In this work we explore Unsupervised Domain Adaptation (UDA) of pretrained language models for downstream tasks. We introduce UDALM, a fine-tuning procedure, using a mixed classification and Masked Language Model loss, t…

Domain AdaptationLanguage ModelingLanguage ModellingSentiment Analysis+1

CONFLATOR: Incorporating Switching Point based Rotatory Positional Encodings for Code-Mixed Language Modeling

2023-09-11 · Mohsin Ali, Kandukuri Sai Teja, Neeharika Gupta, Parth Patwa 외

The mixing of two or more languages is called Code-Mixing (CM). CM is a social norm in multilingual societies. Neural Language Models (NLMs) like transformers have been effective on many NLP tasks. However, NLM for CM is…

Language ModelingLanguage ModellingMachine TranslationSentiment Analysis

MixedPEFT: Combining Multiple PEFT Methods with Mixed Objectives for Unsupervised Domain Adaptation

2026-06-20 · Mohammed Rawhani, Dervis Karaboga, Ozkan Ufuk Nalbantoglu, Alper Basturk 외 arxiv

Pre-trained language models struggle when applied to new domains, as full fine-tuning is computationally expensive and prone to catastrophic forgetting. This study addresses this challenge by presenting a novel parameter…

Unsupervised Domain AdaptationNatural Language Inference

DEMix Layers: Disentangling Domains for Modular Language Modeling

2021-10-16 · ACL ARR October 2021 10 · Anonymous

We introduce a new domain expert mixture (DEMix) layer that enables conditioning a language model (LM) on the domain of the input text. A DEMix layer is a collection of expert feedforward networks, each specialized to a…

Language ModelingLanguage Modelling

DA-Mamba: Domain Adaptive Hybrid Mamba-Transformer Based One-Stage Object Detection

2025-02-16 · A. Enes Doruk, Hasan F. Ates

Recent 2D CNN-based domain adaptation approaches struggle with long-range dependencies due to limited receptive fields, making it difficult to adapt to target domains with significant spatial distribution changes. While …

Domain AdaptationKnowledge DistillationMambaObject+4