paper-with-me

홈 › Papers

A Survey of Spanish Clinical Language Models

2023-08-04 · Guillem García Subies, Álvaro Barbero Jiménez, Paloma Martínez Fernández

This survey focuses in encoder Language Models for solving tasks in the clinical domain in the Spanish language. We review the contributions of 17 corpora focused mainly in clinical tasks, then list the most relevant Spanish Language Models and Spanish Clinical Language models. We perform a thorough comparison of these models by benchmarking them over a curated subset of the available corpora, in order to find the best-performing ones; in total more than 3000 models were fine-tuned for this study. All the tested corpora and the best models are made publically available in an accessible way, so that the results can be reproduced by independent teams or challenged in the future when new Spanish Clinical Language models are created.

📄 PDF Abstract BibTeX arXiv:2308.02199

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingSurvey

Similar Papers 제목 키워드 기반

EriBERTa: A Bilingual Pre-Trained Language Model for Clinical Natural Language Processing

2023-06-12 · Iker de la Iglesia, Aitziber Atutxa, Koldo Gojenola, Ander Barrena

The utilization of clinical reports for various secondary purposes, including health research and treatment monitoring, is crucial for enhancing patient care. Natural Language Processing (NLP) tools have emerged as valua…

Language ModelingLanguage ModellingTransfer Learning

Pretrained Biomedical Language Models for Clinical NLP in Spanish

2022-05-01 · BioNLP (ACL) 2022 5 · Casimiro Pio Carrino, Joan Llop, Marc Pàmies, Asier Gutiérrez-Fandiño 외

This work presents the first large-scale biomedical Spanish language models trained from scratch, using large biomedical corpora consisting of a total of 1.1B tokens and an EHR corpus of 95M tokens. We compared them agai…

NER

Biomedical and Clinical Language Models for Spanish: On the Benefits of Domain-Specific Pretraining in a Mid-Resource Scenario

2021-09-08 · Casimiro Pio Carrino, Jordi Armengol-Estapé, Asier Gutiérrez-Fandiño, Joan Llop-Palao 외

This work presents biomedical and clinical language models for Spanish by experimenting with different pretraining choices, such as masking at word and subword level, varying the vocabulary size and testing with domain d…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

Annotation of negation in the IULA Spanish Clinical Record Corpus

2017-04-01 · WS 2017 4 · Montserrat Marimon, Jorge Vivaldi, N{\'u}ria Bel

This paper presents the IULA Spanish Clinical Record Corpus, a corpus of 3,194 sentences extracted from anonymized clinical records and manually annotated with negation markers and their scope. The corpus was conceived a…

Medical DiagnosisNegationNegation DetectionTerm Extraction

ClinText-SP and RigoBERTa Clinical: a new set of open resources for Spanish Clinical NLP

2025-03-24 · Guillem García Subies, Álvaro Barbero Jiménez, Paloma Martínez Fernández

We present a novel contribution to Spanish clinical natural language processing by introducing the largest publicly available clinical corpus, ClinText-SP, along with a state-of-the-art clinical encoder language model, R…

Language ModelingLanguage Modelling