Rapid Adaptation of BERT for Information Extraction on Domain-Specific Business Documents
Techniques for automatically extracting important content elements from business documents such as contracts, statements, and filings have the potential to make business operations more efficient. This problem can be formulated as a sequence labeling task, and we demonstrate the adaption of BERT to two types of business documents: regulatory filings and property lease agreements. There are aspects of this problem that make it easier than "standard" information extraction tasks and other aspects that make it more difficult, but on balance we find that modest amounts of annotated data (less than 100 documents) are sufficient to achieve reasonable accuracy. We integrate our models into an end-to-end cloud platform that provides both an easy-to-use annotation interface as well as an inference interface that allows users to upload documents and inspect model outputs.
Code (1)
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
DILBERT: Customized Pre-Training for Domain Adaptation withCategory Shift, with an Application to Aspect Extraction
The rise of pre-trained language models has yielded substantial progress in the vast majority of Natural Language Processing (NLP) tasks. However, a generic approach towards the pre-training procedure can naturally be su…
Aspect ExtractionDomain AdaptationLanguage ModellingUnsupervised Domain AdaptationDILBERT: Customized Pre-Training for Domain Adaptation with Category Shift, with an Application to Aspect Extraction
The rise of pre-trained language models has yielded substantial progress in the vast majority of Natural Language Processing (NLP) tasks. However, a generic approach towards the pre-training procedure can naturally be su…
Aspect ExtractionDomain AdaptationLanguage ModellingUnsupervised Domain AdaptationPROTEST-ER: Retraining BERT for Protest Event Extraction
We analyze the effect of further retraining BERT with different domain specific data as an unsupervised domain adaptation strategy for event extraction. Portability of event extraction models is particularly challenging,…
Domain AdaptationEvent ExtractionUnsupervised Domain AdaptationSyntactically Aware Cross-Domain Aspect and Opinion Terms Extraction
A fundamental task of fine-grained sentiment analysis is aspect and opinion terms extraction. Supervised-learning approaches have shown good results for this task; however, they fail to scale across domains where labeled…
Domain AdaptationSentiment AnalysisUnsupervised Domain AdaptationRapid Adaptation of POS Tagging for Domain Specific Uses
Part-of-speech (POS) tagging is a fundamental component for performing natural language tasks such as parsing, information extraction, and question answering. When POS taggers are trained in one domain and applied in sig…
Part-Of-Speech TaggingPOSPOS TaggingQuestion Answering