paper-with-me

홈 › Papers

HistNERo: Historical Named Entity Recognition for the Romanian Language

2024-04-30 · Andrei-Marius Avram, Andreea Iuga, George-Vlad Manolache, Vlad-Cristian Matei, Răzvan-Gabriel Micliuş, Vlad-Andrei Muntean, Manuel-Petru Sorlescu, Dragoş-Andrei Şerban, Adrian-Dinu Urse, Vasile Păiş, Dumitru-Clementin Cercel

This work introduces HistNERo, the first Romanian corpus for Named Entity Recognition (NER) in historical newspapers. The dataset contains 323k tokens of text, covering more than half of the 19th century (i.e., 1817) until the late part of the 20th century (i.e., 1990). Eight native Romanian speakers annotated the dataset with five named entities. The samples belong to one of the following four historical regions of Romania, namely Bessarabia, Moldavia, Transylvania, and Wallachia. We employed this proposed dataset to perform several experiments for NER using Romanian pre-trained language models. Our results show that the best model achieved a strict F1-score of 55.69%. Also, by reducing the discrepancies between regions through a novel domain adaption technique, we improved the performance on this corpus to a strict F1-score of 66.80%, representing an absolute gain of more than 10%.

📄 PDF Abstract BibTeX arXiv:2405.00155

Code (0)

등록된 구현이 없습니다.

Tasks

Domain Adaptationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

Similar Papers 제목 키워드 기반

Introducing RONEC -- the Romanian Named Entity Corpus

2019-09-03 · Stefan Daniel Dumitrescu, Andrei-Marius Avram

We present RONEC - the Named Entity Corpus for the Romanian language. The corpus contains over 26000 entities in ~5000 annotated sentences, belonging to 16 distinct classes. The sentences have been extracted from a copy-…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

Introducing RONEC - the Romanian Named Entity Corpus

2020-05-01 · LREC 2020 5 · Stefan Daniel Dumitrescu, Andrei-Marius Avram

We present RONEC - the Named Entity Corpus for the Romanian language. The corpus contains over 26000 entities in {\textasciitilde}5000 annotated sentences, belonging to 16 distinct classes. The sentences have been extrac…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

Named Entity Recognition in the Romanian Legal Domain

2021-11-01 · EMNLP (NLLP) 2021 11 · Vasile Pais, Maria Mitrofan, Carol Luca Gasan, Vlad Coneschi 외

Recognition of named entities present in text is an important step towards information extraction and natural language understanding. This work presents a named entity recognition system for the Romanian legal domain. Th…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Natural Language Understanding+1

Romanian micro-blogging named entity recognition including health-related entities

2022-10-01 · SMM4H (COLING) 2022 10 · Vasile Pais, Verginica Barbu Mititelu, Elena Irimia, Maria Mitrofan 외

This paper introduces a manually annotated dataset for named entity recognition (NER) in micro-blogging text for Romanian language. It contains gold annotations for 9 entity classes and expressions: persons, locations, o…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

Bootstrapping a Romanian Corpus for Medical Named Entity Recognition

2017-09-01 · RANLP 2017 9 · Maria Mitrofan

Named Entity Recognition (NER) is an important component of natural language processing (NLP), with applicability in biomedical domain, enabling knowledge-discovery from medical texts. Due to the fact that for the Romani…

Medical Named Entity Recognitionnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+3