Domain-Specific Pretraining of Language Models: A Comparative Study in the Medical Field
There are many cases where LLMs are used for specific tasks in a single domain. These usually require less general, but more domain-specific knowledge. Highly capable, general-purpose state-of-the-art language models like GPT-4 or Claude-3-opus can often be used for such tasks, but they are very large and cannot be run locally, even if they were not proprietary. This can be a problem when working with sensitive data. This paper focuses on domain-specific and mixed-domain pretraining as potentially more efficient methods than general pretraining for specialized language models. We will take a look at work related to domain-specific pretraining, specifically in the medical area, and compare benchmark results of specialized language models to general-purpose language models.
Code (0)
등록된 구현이 없습니다.
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A Comparative Study of Pretrained Language Models on Thai Social Text Categorization
The ever-growing volume of data of user-generated content on social media provides a nearly unlimited corpus of unlabeled data even in languages where resources are scarce. In this paper, we demonstrate that state-of-the…
General ClassificationLanguage ModelingLanguage ModellingText CategorizationComparative Study of Language Models on Cross-Domain Data with Model Agnostic Explainability
With the recent influx of bidirectional contextualized transformer language models in the NLP, it becomes a necessity to have a systematic comparative study of these models on variety of datasets. Also, the performance o…
Language ModelingLanguage ModellingComparative Study of Pre-Trained BERT and Large Language Models for Code-Mixed Named Entity Recognition
Named Entity Recognition (NER) in code-mixed text, particularly Hindi-English (Hinglish), presents unique challenges due to informal structure, transliteration, and frequent language switching. This study conducts a comp…
Transfer Learning or Self-supervised Learning? A Tale of Two Pretraining Paradigms
Pretraining has become a standard technique in computer vision and natural language processing, which usually helps to improve performance substantially. Previously, the most dominant pretraining method is transfer learn…
Self-Supervised LearningTransfer LearningSelf-supervised visual learning in the low-data regime: a comparative evaluation
Self-Supervised Learning (SSL) is a valuable and robust training methodology for contemporary Deep Neural Networks (DNNs), enabling unsupervised pretraining on a 'pretext task' that does not require ground-truth labels/a…
Representation LearningSelf-Supervised LearningTransfer Learning