paper-with-me

Papers

Collaboratively Annotating Multilingual Parallel Corpora in the Biomedical Domain---some MANTRAs

2014-05-01 · LREC 2014 5 · Johannes Hellrich, Simon Clematide, Udo Hahn, Dietrich Rebholz-Schuhmann

The coverage of multilingual biomedical resources is high for the English language, yet sparse for non-English languages―an observation which holds for seemingly well-resourced, yet still dramatically low-resourced ones such as Spanish, French or German but even more so for really under-resourced ones such as Dutch. We here present experimental results for automatically annotating parallel corpora and simultaneously acquiring new biomedical terminology for these under-resourced non-English languages on the basis of two types of language resources, namely parallel corpora (i.e. full translation equivalents at the document unit level) and (admittedly deficient) multilingual biomedical terminologies, with English as their anchor language. We automatically annotate these parallel corpora with biomedical named entities by an ensemble of named entity taggers and harmonize non-identical annotations the outcome of which is a so-called silver standard corpus. We conclude with an empirical assessment of this approach to automatically identify both known and new terms in multilingual corpora.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Named Entity Recognition (NER)Translation

Similar Papers 제목 키워드 기반

KBioXLM: A Knowledge-anchored Biomedical Multilingual Pretrained Language Model

2023-11-20 · Lei Geng, Xu Yan, Ziqiang Cao, Juntao Li 외

Most biomedical pretrained language models are monolingual and cannot handle the growing cross-lingual requirements. The scarcity of non-English domain corpora, not to mention parallel data, poses a significant hurdle in…

Language ModelingLanguage ModellingRelationRelation Prediction+1

BVS Corpus: A Multilingual Parallel Corpus of Biomedical Scientific Texts

2019-05-05 · Felipe Soares, Martin Krallinger

The BVS database (Health Virtual Library) is a centralized source of biomedical information for Latin America and Carib, created in 1998 and coordinated by BIREME (Biblioteca Regional de Medicina) in agreement with the P…

ArticlesMachine TranslationSentenceTranslation

FJWU Participation for the WMT21 Biomedical Translation Task

2021-11-01 · WMT (EMNLP) 2021 11 · Sumbal Naz, Sadaf Abdul Rauf, Sami Ul Haq

In this paper we present the FJWU’s system submitted to the biomedical shared task at WMT21. We prepared state-of-the-art multilingual neural machine translation systems for three languages (i.e. German, Spanish and Fren…

Domain AdaptationInformation RetrievalMachine TranslationNMT+2

Building The Sense-Tagged Multilingual Parallel Corpus

2014-05-01 · LREC 2014 5 · Shan Wang, Francis Bond

Sense-annotated parallel corpora play a crucial role in natural language processing. This paper introduces our progress in creating such a corpus for Asian languages using English as a pivot, which is the first such corp…

Biomedical term normalization of EHRs with UMLS

2018-02-08 · LREC 2018 5 · Naiara Perez, Montse Cuadros, German Rigau

This paper presents a novel prototype for biomedical term normalization of electronic health record excerpts with the Unified Medical Language System (UMLS) Metathesaurus. Despite being multilingual and cross-lingual by …