Clinical named entity recognition in the Portuguese language: a benchmark of modern BERT models and LLMs
Clinical notes contain valuable unstructured information. Named entity recognition (NER) enables the automatic extraction of medical concepts; however, benchmarks for Portuguese remain scarce. In this study, we aimed to evaluate BERT-based models and large language models (LLMs) for clinical NER in Portuguese and to test strategies for addressing multilabel imbalance. We compared BioBERTpt, BERTimbau, ModernBERT, and mmBERT with LLMs such as GPT-5 and Gemini-2.5, using the public SemClinBr corpus and a private breast cancer dataset. Models were trained under identical conditions and evaluated using precision, recall, and F1-score. Iterative stratification, weighted loss, and oversampling were explored to mitigate class imbalance. The mmBERT-base model achieved the best performance (micro F1 = 0.76), outperforming all other models. Iterative stratification improved class balance and overall performance. Multilingual BERT models, particularly mmBERT, perform strongly for Portuguese clinical NER and can run locally with limited computational resources. Balanced data-splitting strategies further enhance performance.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Contributions to Clinical Named Entity Recognition in Portuguese
Having in mind that different languages might present different challenges, this paper presents the following contributions to the area of Information Extraction from clinical text, targeting the Portuguese language: a c…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Word EmbeddingsBioBERTpt - A Portuguese Neural Language Model for Clinical Named Entity Recognition
With the growing number of electronic health record data, clinical NLP tasks have become increasingly relevant to unlock valuable information from unstructured clinical text. Although the performance of downstream NLP ta…
Language ModelingLanguage Modellingnamed-entity-recognitionNamed Entity Recognition+3PAMPO: using pattern matching and pos-tagging for effective Named Entities recognition in Portuguese
This paper deals with the entity extraction task (named entity recognition) of a text mining process that aims at unveiling non-trivial semantic structures, such as relationships and interaction between entities or commu…
Entity Extraction using GANnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+4HAREM: the first evaluation contest for Named Entity Recognition in Portuguese
The first evaluation contest for Named Entity Recognition in Portuguese
named-entity-recognitionNamed Entity RecognitionNER