paper-with-me

Papers

Spanish Biomedical Crawled Corpus: A Large, Diverse Dataset for Spanish Biomedical Language Models

2021-09-16 · Casimiro Pio Carrino, Jordi Armengol-Estapé, Ona de Gibert Bonet, Asier Gutiérrez-Fandiño, Aitor Gonzalez-Agirre, Martin Krallinger, Marta Villegas

We introduce CoWeSe (the Corpus Web Salud Espa\~nol), the largest Spanish biomedical corpus to date, consisting of 4.5GB (about 750M tokens) of clean plain text. CoWeSe is the result of a massive crawler on 3000 Spanish domains executed in 2020. The corpus is openly available and already preprocessed. CoWeSe is an important resource for biomedical and health NLP in Spanish and has already been employed to train domain-specific language models and to produce word embbedings. We released the CoWeSe corpus under a Creative Commons Attribution 4.0 International license, both in Zenodo (\url{https://zenodo.org/record/4561971\#.YTI5SnVKiEA}).

📄 PDF Abstract BibTeX arXiv:2109.07765

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MarIA: Spanish Language Models

2021-07-15 · Asier Gutiérrez-Fandiño, Jordi Armengol-Estapé, Marc Pàmies, Joan Llop-Palao 외

This work presents MarIA, a family of Spanish language models and associated resources made available to the industry and the research community. Currently, MarIA includes RoBERTa-base, RoBERTa-large, GPT2 and GPT2-large…

Extractive Question-AnsweringQuestion Answering

NUBES: A Corpus of Negation and Uncertainty in Spanish Clinical Texts

2020-04-02 · LREC 2020 5 · Salvador Lima, Naiara Perez, Montse Cuadros, German Rigau

This paper introduces the first version of the NUBes corpus (Negation and Uncertainty annotations in Biomedical texts in Spanish). The corpus is part of an on-going research and currently consists of 29,682 sentences obt…

Negation

Pretrained Biomedical Language Models for Clinical NLP in Spanish

2022-05-01 · BioNLP (ACL) 2022 5 · Casimiro Pio Carrino, Joan Llop, Marc Pàmies, Asier Gutiérrez-Fandiño 외

This work presents the first large-scale biomedical Spanish language models trained from scratch, using large biomedical corpora consisting of a total of 1.1B tokens and an EHR corpus of 95M tokens. We compared them agai…

NER

When Specialization Helps: Using Pooled Contextualized Embeddings to Detect Chemical and Biomedical Entities in Spanish

2019-10-08 · WS 2019 11 · Manuel Stoeckel, Wahed Hemati, Alexander Mehler

The recognition of pharmacological substances, compounds and proteins is an essential preliminary work for the recognition of relations between chemicals and other biomedically relevant units. In this paper, we describe …

ArticlesWord Embeddings

BVS Corpus: A Multilingual Parallel Corpus of Biomedical Scientific Texts

2019-05-05 · Felipe Soares, Martin Krallinger

The BVS database (Health Virtual Library) is a centralized source of biomedical information for Latin America and Carib, created in 1998 and coordinated by BIREME (Biblioteca Regional de Medicina) in agreement with the P…

ArticlesMachine TranslationSentenceTranslation