paper-with-me

Papers

Local LLM Ensembles for Zero-shot Portuguese Named Entity Recognition

2025-12-10 · João Lucas Luz Lima Sarcinelli, Diego Furtado Silva arxiv

Large Language Models (LLMs) excel in many Natural Language Processing (NLP) tasks through in-context learning but often under-perform in Named Entity Recognition (NER), especially for lower-resource languages like Portuguese. While open-weight LLMs enable local deployment, no single model dominates all tasks, motivating ensemble approaches. However, existing LLM ensembles focus on text generation or classification, leaving NER under-explored. In this context, this work proposes a novel three-step ensemble pipeline for zero-shot NER using similarly capable, locally run LLMs. Our method outperforms individual LLMs in four out of five Portuguese NER datasets by leveraging a heuristic to select optimal model combinations with minimal annotated data. Moreover, we show that ensembles obtained on different source datasets generally outperform individual LLMs in cross-dataset configurations, potentially eliminating the need for annotated data for the current task. Our work advances scalable, low-resource, and zero-shot NER by effectively combining multiple small LLMs without fine-tuning. Code is available at https://github.com/Joao-Luz/local-llm-ner-ensemble.

📄 PDF Abstract BibTeX arXiv:2512.10043

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

Introducing Bode: A Fine-Tuned Large Language Model for Portuguese Prompt-Based Task

2024-01-05 · Gabriel Lino Garcia, Pedro Henrique Paiola, Luis Henrique Morelli, Giovani Candido 외

Large Language Models (LLMs) are increasingly bringing advances to Natural Language Processing. However, low-resource languages, those lacking extensive prominence in datasets for various NLP tasks, or where existing dat…

In-Context LearningLanguage ModelingLanguage ModellingLarge Language Model

CAMÕES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese

2025-08-27 · Carlos Carvalho, Francisco Teixeira, Catarina Botelho, Anna Pompili 외 arxiv

Existing resources for Automatic Speech Recognition in Portuguese are mostly focused on Brazilian Portuguese, leaving European Portuguese (EP) and other varieties under-explored. To bridge this gap, we introduce CAMÕES, …

Speech Recognition

Portuguese Named Entity Recognition using Conditional Random Fields and Local Grammars

2018-05-01 · LREC 2018 5 · Juliana Pirovani, Elias Oliveira
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Question Answering+1

Can Large Language Models Reliably Extract Physiology Index Values from Coronary Angiography Reports?

2026-03-21 · Sofia Morgado, Filipa Valdeira, Niklas Sander, Diogo Ferreira 외 arxiv

Coronary angiography (CAG) reports contain clinically relevant physiological measurements, yet this information is typically in the form of unstructured natural language, limiting its use in research. We investigate the …

HAREM: the first evaluation contest for Named Entity Recognition in Portuguese

2006-02-24 · - 2006 2 · Diana Santos

The first evaluation contest for Named Entity Recognition in Portuguese

named-entity-recognitionNamed Entity RecognitionNER