paper-with-me

홈 › Papers

Customizing Contextualized Language Models forLegal Document Reviews

2021-02-10 · Shohreh Shaghaghian, Luna, Feng, Borna Jafarpour, Nicolai Pogrebnyakov

Inspired by the inductive transfer learning on computer vision, many efforts have been made to train contextualized language models that boost the performance of natural language processing tasks. These models are mostly trained on large general-domain corpora such as news, books, or Wikipedia.Although these pre-trained generic language models well perceive the semantic and syntactic essence of a language structure, exploiting them in a real-world domain-specific scenario still needs some practical considerations to be taken into account such as token distribution shifts, inference time, memory, and their simultaneous proficiency in multiple tasks. In this paper, we focus on the legal domain and present how different language model strained on general-domain corpora can be best customized for multiple legal document reviewing tasks. We compare their efficiencies with respect to task performances and present practical considerations.

📄 PDF Abstract BibTeX arXiv:2102.05757

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingTransfer Learning

Similar Papers 제목 키워드 기반

Cross-lingual Contextualized Topic Models with Zero-shot Learning

2020-04-16 · EACL 2021 2 · Federico Bianchi, Silvia Terragni, Dirk Hovy, Debora Nozza 외

Many data sets (e.g., reviews, forums, news, etc.) exist parallelly in multiple languages. They all cover the same content, but the linguistic differences make it impossible to use traditional, bag-of-word-based topic mo…

Topic ModelsTransfer LearningVariational InferenceZero-Shot Learning

Says Who\ldots? Identification of Expert versus Layman Critics' Reviews of Documentary Films

2016-12-01 · COLING 2016 12 · Ming Jiang, Jana Diesner

We extend classic review mining work by building a binary classifier that predicts whether a review of a documentary film was written by an expert or a layman with 90.70{\%} accuracy (F1 score), and compare the character…

Decision MakingDiversityRecommendation Systems

CEDR: Contextualized Embeddings for Document Ranking

2019-04-15 · Sean MacAvaney, Andrew Yates, Arman Cohan, Nazli Goharian

Although considerable attention has been given to neural ranking architectures recently, far less attention has been paid to the term representations that are used as input to these models. In this work, we investigate h…

Ad-Hoc Information RetrievalDocument RankingGeneral Classification

BERTgrid: Contextualized Embedding for 2D Document Representation and Understanding

2019-09-11 · NeurIPS Workshop Document_Intelligen 2019 12 · Timo I. Denk, Christian Reisswig

For understanding generic documents, information like font sizes, column layout, and generally the positioning of words may carry semantic information that is crucial for solving a downstream document intelligence task. …

Instance SegmentationLanguage ModelingLanguage ModellingSemantic Segmentation

Improving Contextualized Topic Models with Negative Sampling

2023-03-27 · Suman Adhya, Avishek Lahiri, Debarshi Kumar Sanyal, Partha Pratim Das

Topic modeling has emerged as a dominant method for exploring large document collections. Recent approaches to topic modeling use large contextualized language models and variational autoencoders. In this paper, we propo…

DiversityTopic ModelsTriplet