paper-with-me

홈 › Papers

The Russian Legislative Corpus

2024-06-07 · Denis Saveliev, Ruslan Kuchakov

We present the comprehensive Russian primary and secondary legislation corpus covering 1991 to 2023. The corpus collects all 281,413 texts (176,523,268 tokens) of non-secret federal regulations and acts, along with their metadata. The corpus has two versions the original text with minimal preprocessing and a version prepared for linguistic analysis with morphosyntactic markup.

📄 PDF Abstract BibTeX arXiv:2406.04855

Code (1)

irlcode/RusLawOD 공식 구현

Similar Papers 제목 키워드 기반

UlyssesNER-Br: A Corpus of Brazilian Legislative Documents for Named Entity Recognition

2022-03-16 · PROPOR 2022 2022 3 · Hidelberg O. Albuquerque, Rosimeire Costa, Gabriel Silvestre, Ellen Souza 외

The amount of legislative documents produced within the past decade has risen dramatically, making it difficult for law practitioners to consult and update legislation. Named Entity Recognition (NER) systems have the unt…

Decision MakingInformation Retrievalnamed-entity-recognitionNamed Entity Recognition+3

Collection and Annotation of the Romanian Legal Corpus

2020-05-01 · LREC 2020 5 · Dan Tufi{\textcommabelow{s}}, Maria Mitrofan, Vasile P{\u{a}}i{\textcommabelow{s}}, Radu Ion 외

We present the Romanian legislative corpus which is a valuable linguistic asset for the development of machine translation systems, especially for under-resourced languages. The knowledge that can be extracted from this …

Machine TranslationPOSTranslation

Sense-Annotated Corpus for Russian

2022-09-01 · CLIB 2022 9 · Alexander Kirillovich, Natalia Loukachevitch, Maksim Kulaev, Angelina Bolshina 외

We present a sense-annotated corpus for Russian. The resource was obtained my manually annotating texts from the OpenCorpora corpus, an open corpus for the Russian language, by senses of Russian wordnet RuWordNet. The an…

Word Sense Disambiguation

Event2Mind for Russian: Understanding Emotions and Intents in Texts. Corpus and Model for Evaluation

2020-06-17 · Computational Linguistics and Intellectual Technologies: Proceedings of the International Conference “Dialogue 2020” 2020 6 · Fenogenova A. S., Tikhonova M. I., Filipetskaya D. V., Mironenko F. D. 외

The paper provides a comprehensive overview of the corpus for the Russian language for the commonsense inference task. Namely, we construct event phrases, which cover a wide range of everyday situations with labelled int…

Common Sense Reasoning

Categorisation of Bulgarian Legislative Documents

2020-09-01 · CLIB 2020 9 · Nikola Obreshkov, Martin Yalamov, Svetla Koeva

The paper presents the categorisation of Bulgarian MARCELL corpus in toplevel EuroVoc domains. The Bulgarian MARCELL corpus is part of a recently developed multilingual corpus representing the national legislation in sev…

Term Extraction