paper-with-me

Papers

KIND: an Italian Multi-Domain Dataset for Named Entity Recognition

2021-12-30 · LREC 2022 6 · Teresa Paccosi, Alessio Palmero Aprosio

In this paper we present KIND, an Italian dataset for Named-entity recognition. It contains more than one million tokens with annotation covering three classes: person, location, and organization. The dataset (around 600K tokens) mostly contains manual gold annotations in three different domains (news, literature, and political discourses) and a semi-automatically annotated part. The multi-domain feature is the main strength of the present work, offering a resource which covers different styles and language uses, as well as the largest Italian NER dataset with manual gold annotations. It represents an important resource for the training of NER systems in Italian. Texts and annotations are freely downloadable from the Github repository.

📄 PDF Abstract BibTeX arXiv:2112.15099

Code (1)

dhfbk/kind 공식 구현

Tasks

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

Similar Papers 제목 키워드 기반

ENEIDE: A High Quality Silver Standard Dataset for Named Entity Recognition and Linking in Historical Italian

2026-03-31 · Cristian Santini, Sebastian Barzaghi, Paolo Sernani, Emanuele Frontoni 외 arxiv

This paper introduces ENEIDE (Extracting Named Entities from Italian Digital Editions), a silver standard dataset for Named Entity Recognition and Linking (NERL) in historical Italian texts. The corpus comprises 2,111 do…

Entity Disambiguation

BERTino: an Italian DistilBERT model

2023-03-31 · Matteo Muffo, Enrico Bertino

The recent introduction of Transformers language representation models allowed great improvements in many natural language processing (NLP) tasks. However, if on one hand the performances achieved by this kind of archite…

model

Named Entity Recognition in Historical Italian: The Case of Giacomo Leopardi's Zibaldone

2025-05-26 · Cristian Santini, Laura Melosi, Emanuele Frontoni

The increased digitization of world's textual heritage poses significant challenges for both computer science and literary studies. Overall, there is an urgent need of computational techniques able to adapt to the challe…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

Multilevel Annotation of Agreement and Disagreement in Italian News Blogs

2016-05-01 · LREC 2016 5 · Fabio Celli, Giuseppe Riccardi, Firoj Alam

In this paper, we present a corpus of news blog conversations in Italian annotated with gold standard agreement/disagreement relations at message and sentence levels. This is the first resource of this kind in Italian. F…

Sentence

SLIMER-IT: Zero-Shot NER on Italian Language

2024-09-24 · Andrew Zamai, Leonardo Rigutini, Marco Maggini, Andrea Zugarini

Traditional approaches to Named Entity Recognition (NER) frame the task into a BIO sequence labeling problem. Although these systems often excel in the downstream task at hand, they require extensive annotated data and s…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER