Go Simple and Pre-Train on Domain-Specific Corpora: On the Role of Training Data for Text Classification
Pre-trained language models provide the foundations for state-of-the-art performance across a wide range of natural language processing tasks, including text classification. However, most classification datasets assume a large amount labeled data, which is commonly not the case in practical settings. In particular, in this paper we compare the performance of a light-weight linear classifier based on word embeddings, i.e., fastText (Joulin et al., 2017), versus a pre-trained language model, i.e., BERT (Devlin et al., 2019), across a wide range of datasets and classification tasks. In general, results show the importance of domain-specific unlabeled data, both in the form of word embeddings or language models. As for the comparison, BERT outperforms all baselines in standard datasets with large training sets. However, in settings with small training datasets a simple method like fastText coupled with domain-specific word embeddings performs equally well or better than BERT, even when pre-trained on domain-specific data.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationLanguage ModelingLanguage Modellingtext-classificationText ClassificationWord EmbeddingsMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Domain Adaptation for NMT via Filtered Iterative Back-Translation
Domain-specific Neural Machine Translation (NMT) model can provide improved performance, however, it is difficult to always access a domain-specific parallel corpus. Iterative Back-Translation can be used for fine-tuning…
Domain AdaptationMachine TranslationNMTTranslationBilingual Terminology Extraction Using Neural Word Embeddings on Comparable Corpora
Term and glossary management are vital steps of preparation of every language specialist, and they play a very important role at the stage of education of translation professionals. The growing trend of efficient time ma…
ManagementRetrievalTranslationWord EmbeddingsNot just about size - A Study on the Role of Distributed Word Representations in the Analysis of Scientific Publications
The emergence of knowledge graphs in the scholarly communication domain and recent advances in artificial intelligence and natural language processing bring us closer to a scenario where intelligent systems can assist sc…
Knowledge GraphsSpecificityWord EmbeddingsAttention-Driven Multi-Agent Reinforcement Learning: Enhancing Decisions with Expertise-Informed Tasks
In this paper, we introduce an alternative approach to enhancing Multi-Agent Reinforcement Learning (MARL) through the integration of domain knowledge and attention-based policy mechanisms. Our methodology focuses on the…
Decision MakingMulti-agent Reinforcement LearningThe Robotic Surgery Procedural Framebank
Robot-Assisted minimally invasive robotic surgery is the gold standard for the surgical treatment of many pathological conditions, and several manuals and academic papers describe how to perform these interventions. Thes…
Natural Language UnderstandingSemantic ParsingSemantic Role Labeling