paper-with-me

홈 › Papers

BioMedBERT: A Pre-trained Biomedical Language Model for QA and IR

2020-12-10 · S Chakraborty, E Bisong, S Bhatt, T Wagner, R Elliott, F Mosconi

The SARS-CoV-2 (COVID-19) pandemic spotlighted the importance of moving quickly with biomedical research. However, as the number of biomedical research papers continue to increase, the task of finding relevant articles to answer pressing questions has become significant. In this work, we propose a textual data mining tool that supports literature search to accelerate the work of researchers in the biomedical domain. We achieve this by building a neural-based deep contextual understanding model for Question-Answering (QA) and Information Retrieval (IR) tasks. We also leverage the new BREATHE dataset which is one of the largest available datasets of biomedical research literature, containing abstracts and full-text articles from ten different biomedical literature sources on which we pre-train our BioMedBERT model. Our work achieves state-of-the-art results on the QA fine-tuning task on BioASQ 5b, 6b and 7b datasets. In addition, we observe superior relevant results when BioMedBERT embeddings are used with Elasticsearch for the Information Retrieval task on the intelligently formulated BioASQ dataset. We believe our diverse dataset and our unique model architecture are what led us to achieve the state-of-the-art results for QA and IR tasks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesInformation RetrievalLanguage ModelingLanguage ModellingQuestion AnsweringRetrieval

Similar Papers 제목 키워드 기반

Harnessing Large Language Models for Biomedical Named Entity Recognition

2025-12-28 · Jian Chen, Leilei Su, Cong Sun arxiv

Background and Objective: Biomedical Named Entity Recognition (BioNER) is a foundational task in medical informatics, crucial for downstream applications like drug discovery and clinical trial matching. However, adapting…

Drug Discovery

Evaluation design conditions the expert-vs-auto MeSH gap: a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark

2026-07-23 · Samuel M. Okoe-Mensah arxiv

A systematic review begins with someone reading thousands of abstracts to identify the few that are relevant, and classifiers are used to prioritise that reading. Their inputs are often augmented with Medical Subject Hea…

BioBERT: a pre-trained biomedical language representation model for biomedical text mining

2019-01-25 · Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim 외

Biomedical text mining is becoming increasingly important as the number of biomedical documents rapidly grows. With the progress in natural language processing (NLP), extracting valuable information from biomedical liter…

Drug–drug Interaction ExtractionFew-Shot LearningLanguage ModellingMedical Named Entity Recognition+8

Pre-trained Language Models in Biomedical Domain: A Systematic Survey

2021-10-11 · Benyou Wang, Qianqian Xie, Jiahuan Pei, Zhihong Chen 외

Pre-trained language models (PLMs) have been the de facto paradigm for most natural language processing (NLP) tasks. This also benefits biomedical domain: researchers from informatics, medicine, and computer science (CS)…

Survey

BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining

2022-10-19 · Renqian Luo, Liai Sun, Yingce Xia, Tao Qin 외

Pre-trained language models have attracted increasing attention in the biomedical domain, inspired by their great success in the general natural language domain. Among the two main branches of pre-trained language models…

Document ClassificationLanguage ModellingQuestion AnsweringRelation Extraction+1