paper-with-me

홈 › Papers

MedJEx: A Medical Jargon Extraction Model with Wiki's Hyperlink Span and Contextualized Masked Language Model Score

2022-10-12 · Sunjae Kwon, Zonghai Yao, Harmon S. Jordan, David A. Levy, Brian Corner, Hong Yu

This paper proposes a new natural language processing (NLP) application for identifying medical jargon terms potentially difficult for patients to comprehend from electronic health record (EHR) notes. We first present a novel and publicly available dataset with expert-annotated medical jargon terms from 18K+ EHR note sentences ($MedJ$). Then, we introduce a novel medical jargon extraction ($MedJEx$) model which has been shown to outperform existing state-of-the-art NLP models. First, MedJEx improved the overall performance when it was trained on an auxiliary Wikipedia hyperlink span dataset, where hyperlink spans provide additional Wikipedia articles to explain the spans (or terms), and then fine-tuned on the annotated MedJ data. Secondly, we found that a contextualized masked language model score was beneficial for detecting domain-specific unfamiliar jargon terms. Moreover, our results show that training on the auxiliary Wikipedia hyperlink span datasets improved six out of eight biomedical named entity recognition benchmark datasets. Both MedJ and MedJEx are publicly available.

📄 PDF Abstract BibTeX arXiv:2210.05875

Code (1)

mozzitastebitter/medjex 공식 구현 pytorch

Tasks

ArticlesLanguage ModelingLanguage Modellingmodelnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

Similar Papers 제목 키워드 기반

Large Language Model-based Role-Playing for Personalized Medical Jargon Extraction

2024-08-10 · Jung Hoon Lim, Sunjae Kwon, Zonghai Yao, John P. Lalor 외

Previous studies reveal that Electronic Health Records (EHR), which have been widely adopted in the U.S. to allow patients to access their personal medical information, do not have high readability to patients due to the…

In-Context LearningLanguage ModelingLanguage ModellingLarge Language Model+1

HopRetriever: Retrieve Hops over Wikipedia to Answer Complex Questions

2020-12-31 · Shaobo Li, Xiaoguang Li, Lifeng Shang, Xin Jiang 외

Collecting supporting evidence from large corpora of text (e.g., Wikipedia) is of great challenge for open-domain Question Answering (QA). Especially, for multi-hop open-domain QA, scattered evidence pieces are required …

Document EmbeddingOpen-Domain Question AnsweringQuestion AnsweringRetrieval

Automatic Construction of a Large-Scale Corpus for Geoparsing Using Wikipedia Hyperlinks

2024-03-25 · Keyaki Ohno, Hirotaka Kameko, Keisuke Shirai, Taichi Nishimura 외

Geoparsing is the task of estimating the latitude and longitude (coordinates) of location expressions in texts. Geoparsing must deal with the ambiguity of the expressions that indicate multiple locations with the same no…

Articles

WEXEA: Wikipedia EXhaustive Entity Annotation

2020-05-01 · LREC 2020 5 · Michael Strobl, Amine Trabelsi, Osmar Zaiane

Building predictive models for information extraction from text, such as named entity recognition or the extraction of semantic relationships between named entities in text, requires a large corpus of annotated text. Wik…

Articlesnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

Injecting Knowledge Base Information into End-to-End Joint Entity and Relation Extraction and Coreference Resolution

2021-07-05 · Findings (ACL) 2021 8 · Severine Verlinden, Klim Zaporojets, Johannes Deleu, Thomas Demeester 외

We consider a joint information extraction (IE) model, solving named entity recognition, coreference resolution and relation extraction jointly over the whole document. In particular, we study how to inject information f…

coreference-resolutionCoreference ResolutionEntity LinkingJoint Entity and Relation Extraction+4