Evaluating Pre-Trained Language Models for Focused Terminology Extraction from Swedish Medical Records
In the experiments briefly presented in this abstract, we compare the performance of a generalist Swedish pre-trained language model with a domain-specific Swedish pre-trained model on the downstream task of focussed terminology extraction of implant terms, which are terms that indicate the presence of implants in the body of patients. The fine-tuning is identical for both models. For the search strategy we rely on KD-Tree that we feed with two different lists of term seeds, one with noise and one without noise. Results shows that the use of a domain-specific pre-trained language model has a positive impact on focussed terminology extraction only when using term seeds without noise.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
Adapting and evaluating a generic term extraction tool
We present techniques for monolingual term candidate extraction which are being developed in the EU project TTC. We designed an application for German and English data that serves as a first evaluation of the methods for…
LemmatizationTerm ExtractionEntityBERT: Entity-centric Masking Strategy for Model Pretraining for the Clinical Domain
Transformer-based neural language models have led to breakthroughs for a variety of natural language processing (NLP) tasks. However, most models are pretrained on general domain data. We propose a methodology to produce…
NegationNegation DetectionRelationRelation Extraction+1Evaluating the Reliability and Interaction of Recursively Used Feature Classes for Terminology Extraction
Feature design and selection is a crucial aspect when treating terminology extraction as a machine learning classification problem. We designed feature classes which characterize different properties of terms based on di…
BIG-bench Machine LearningClassificationGeneral ClassificationMachine Translation+1An Entity-based Claim Extraction Pipeline for Real-world Biomedical Fact-checking
Existing fact-checking models for biomedical claims are typically trained on synthetic or well-worded data and hardly transfer to social media content. This mismatch can be mitigated by adapting the social media input to…
Entity LinkingFact Checkingnamed-entity-recognitionNamed Entity Recognition+1Bootstrapping Term Extractors for Multiple Languages
Terminology extraction resources are needed for a wide range of human language technology applications, including knowledge management, information extraction, semantic search, cross-language information retrieval and au…
Information RetrievalManagementPOSRetrieval+2