Biomedical Nested NER with Large Language Model and UMLS Heuristics
In this paper, we present our system for the BioNNE English track, which aims to extract 8 types of biomedical nested named entities from biomedical text. We use a large language model (Mixtral 8x7B instruct) and ScispaCy NER model to identify entities in an article and build custom heuristics based on unified medical language system (UMLS) semantic types to categorize the entities. We discuss the results and limitations of our system and propose future improvements. Our system achieved an F1 score of 0.39 on the BioNNE validation set and 0.348 on the test set.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage ModellingLarge Language ModelNERMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Biomedical Event Extraction with Hierarchical Knowledge Graphs
Biomedical event extraction is critical in understanding biomolecular interactions described in scientific corpus. One of the main challenges is to identify nested structured events that are associated with non-indicativ…
Event ExtractionLanguage ModelingSentenceEvaluating Biomedical BERT Models for Vocabulary Alignment at Scale in the UMLS Metathesaurus
The current UMLS (Unified Medical Language System) Metathesaurus construction process for integrating over 200 biomedical source vocabularies is expensive and error-prone as it relies on the lexical algorithms and human …
Task 2Word EmbeddingsInjecting Structured Biomedical Knowledge into Language Models: Continual Pretraining vs. GraphRAG
The injection of domain-specific knowledge is crucial for adapting language models (LMs) to specialized fields such as biomedicine. While most current approaches rely on unstructured text corpora, this study explores two…
Continual PretrainingQuestion AnsweringHILGEN: Hierarchically-Informed Data Generation for Biomedical NER Using Knowledgebases and Large Language Models
We present HILGEN, a Hierarchically-Informed Data Generation approach that combines domain knowledge from the Unified Medical Language System (UMLS) with synthetic data generated by large language models (LLMs), specific…
Data AugmentationNERSynthetic Data GenerationUBERT: A Novel Language Model for Synonymy Prediction at Scale in the UMLS Metathesaurus
The UMLS Metathesaurus integrates more than 200 biomedical source vocabularies. During the Metathesaurus construction process, synonymous terms are clustered into concepts by human editors, assisted by lexical similarity…
Language ModelingLanguage ModellingPredictionSentence