paper-with-me

홈 › Papers

UlyssesNER-Br: A Corpus of Brazilian Legislative Documents for Named Entity Recognition

2022-03-16 · PROPOR 2022 2022 3 · Hidelberg O. Albuquerque, Rosimeire Costa, Gabriel Silvestre, Ellen Souza, Nádia F. F. da Silva, Douglas Vitório, Gyovana Moriyama, Lucas Martins, Luiza Soezima, Augusto Nunes, Felipe Siqueira, João P. Tarrega, Joao V. Beinotti, Marcio Dias, Matheus Silva, Miguel Gardini, Vinicius Silva, André C. P. L. F. de Carvalho, Adriano L. I. Oliveira

The amount of legislative documents produced within the past decade has risen dramatically, making it difficult for law practitioners to consult and update legislation. Named Entity Recognition (NER) systems have the untapped potential to extract information from legal documents, which can improve information retrieval and decision-making processes. We introduce the UlyssesNER-Br, a corpus of Brazilian Legislative Documents for NER with quality baselines. The presented corpus consists of bills and legislative consultations from Brazilian Chamber of Deputies. We implemented Conditional Random Field (CRF) and Hidden Markov Model (HMM) models, and the promising F1-score of 80.8% in the analysis by categories and 81.04% in the analysis by types, was achieved with the CRF model. The entities with the best average F1- score results were “FUNDlei” and “DATA”, and the ones with the worst results were “EVENTO” and “PESSOAgrupoind”. The corpus was also evaluated using a BiLSTM-CRF and Glove architectures provided by the pioneering state-of-the-art paper, achieving F1-score of 76.89% in the analysis by categories and 59.67% in the analysis by types.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingInformation Retrievalnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NERRetrieval

Similar Papers 제목 키워드 기반

Collection and Annotation of the Romanian Legal Corpus

2020-05-01 · LREC 2020 5 · Dan Tufi{\textcommabelow{s}}, Maria Mitrofan, Vasile P{\u{a}}i{\textcommabelow{s}}, Radu Ion 외

We present the Romanian legislative corpus which is a valuable linguistic asset for the development of machine translation systems, especially for under-resourced languages. The knowledge that can be extracted from this …

Machine TranslationPOSTranslation

The MARCELL Legislative Corpus

2020-05-01 · LREC 2020 5 · Tam{\'a}s V{\'a}radi, Svetla Koeva, Martin Yamalov, Marko Tadi{\'c} 외

This article presents the current outcomes of the MARCELL CEF Telecom project aiming to collect and deeply annotate a large comparable corpus of legal documents. The MARCELL corpus includes 7 monolingual sub-corpora (Bul…

Sentence

SinhaLegal: A Benchmark Corpus for Information Extraction and Analysis in Sinhala Legislative Texts

2026-03-05 · Minduli Lasandi, Nevidu Jayatilleke arxiv

SinhaLegal introduces a Sinhala legislative text corpus containing approximately 2 million words across 1,206 legal documents. The dataset includes two types of legal documents: 1,065 Acts dated from 1981 to 2014 and 141…

Information ExtractionDocument AI

Categorisation of Bulgarian Legislative Documents

2020-09-01 · CLIB 2020 9 · Nikola Obreshkov, Martin Yalamov, Svetla Koeva

The paper presents the categorisation of Bulgarian MARCELL corpus in toplevel EuroVoc domains. The Bulgarian MARCELL corpus is part of a recently developed multilingual corpus representing the national legislation in sev…

Term Extraction

Natural Language Processing Pipeline to Annotate Bulgarian Legislative Documents

2020-05-01 · LREC 2020 5 · Svetla Koeva, Nikola Obreshkov, Martin Yalamov

The paper presents the Bulgarian MARCELL corpus, part of a recently developed multilingual corpus representing the national legislation in seven European countries and the NLP pipeline that turns the web crawled data int…

Sentence