Towards a comprehensive open repository of Polish language resources
The aim of this paper is to present current efforts towards the creation of a comprehensive open repository of Polish language resources and tools (LRTs). The work described here is carried out within the CESAR project, member of the META-NET consortium. It has already resulted in the creation of the Computational Linguistics in Poland site containing an exhaustive collection of Polish LRTs. Current work is focused on the creation of new LRTs and, esp., the enhancement of existing LRTs, such as parallel corpora, annotated corpora of written and spoken Polish and morphological dictionaries to be made available via the META-SHARE repository.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Open Repository of the Polish Sign Language Corpus: Publication Project of the Polish Sign Language Corpus
Between 2010 and 2020, the research team of the Section for Sign Linguistics collected, annotated, and translated a large corpus of Polish Sign Language (polski język migowy, PJM). After this task was finished, a substan…
Towards Mapping Thesauri onto plWordNet
plWordNet, the wordnet of Polish, has become a very comprehensive description of the Polish lexical system. This paper presents a plan of its semi-automated integration with thesauri, terminological databases and ontolog…
Keyword ExtractionInforex --- a collaborative system for text corpora annotation and analysis
We report a first major upgrade of Inforex {---} a web-based system for qualitative and collaborative text corpora annotation and analysis. Inforex is a part of Polish CLARIN infrastructure. It is integrated with a digit…
Named Entity Recognition (NER)Word Sense DisambiguationLanguageCrawl: A Generic Tool for Building Language Models Upon Common-Crawl
The web data contains immense amount of data, hundreds of billion words are waiting to be extracted and used for language research. In this work we introduce our tool LanguageCrawl which allows NLP researchers to easily …
Language ModelingLanguage ModellingBAN-PL: a Novel Polish Dataset of Banned Harmful and Offensive Content from Wykop.pl web service
Since the Internet is flooded with hate, it is one of the main tasks for NLP experts to master automated online content moderation. However, advancements in this field require improved access to publicly available accura…
Specificity