The Russian Legislative Corpus
We present the comprehensive Russian primary and secondary legislation corpus covering 1991 to 2023. The corpus collects all 281,413 texts (176,523,268 tokens) of non-secret federal regulations and acts, along with their metadata. The corpus has two versions the original text with minimal preprocessing and a version prepared for linguistic analysis with morphosyntactic markup.
Code (1)
Similar Papers 제목 키워드 기반
UlyssesNER-Br: A Corpus of Brazilian Legislative Documents for Named Entity Recognition
The amount of legislative documents produced within the past decade has risen dramatically, making it difficult for law practitioners to consult and update legislation. Named Entity Recognition (NER) systems have the unt…
Decision MakingInformation Retrievalnamed-entity-recognitionNamed Entity Recognition+3Collection and Annotation of the Romanian Legal Corpus
We present the Romanian legislative corpus which is a valuable linguistic asset for the development of machine translation systems, especially for under-resourced languages. The knowledge that can be extracted from this …
Machine TranslationPOSTranslationSense-Annotated Corpus for Russian
We present a sense-annotated corpus for Russian. The resource was obtained my manually annotating texts from the OpenCorpora corpus, an open corpus for the Russian language, by senses of Russian wordnet RuWordNet. The an…
Word Sense DisambiguationEvent2Mind for Russian: Understanding Emotions and Intents in Texts. Corpus and Model for Evaluation
The paper provides a comprehensive overview of the corpus for the Russian language for the commonsense inference task. Namely, we construct event phrases, which cover a wide range of everyday situations with labelled int…
Common Sense ReasoningCategorisation of Bulgarian Legislative Documents
The paper presents the categorisation of Bulgarian MARCELL corpus in toplevel EuroVoc domains. The Bulgarian MARCELL corpus is part of a recently developed multilingual corpus representing the national legislation in sev…
Term Extraction