paper-with-me

홈 › Papers

Towards a comprehensive open repository of Polish language resources

2012-05-01 · LREC 2012 5 · Maciej Ogrodniczuk, Piotr P{\k{e}}zik, Adam Przepi{\'o}rkowski

The aim of this paper is to present current efforts towards the creation of a comprehensive open repository of Polish language resources and tools (LRTs). The work described here is carried out within the CESAR project, member of the META-NET consortium. It has already resulted in the creation of the Computational Linguistics in Poland site containing an exhaustive collection of Polish LRTs. Current work is focused on the creation of new LRTs and, esp., the enhancement of existing LRTs, such as parallel corpora, annotated corpora of written and spoken Polish and morphological dictionaries to be made available via the META-SHARE repository.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Open Repository of the Polish Sign Language Corpus: Publication Project of the Polish Sign Language Corpus

2022-06-01 · SignLang (LREC) 2022 6 · Anna Kuder, Joanna Wójcicka, Piotr Mostowski, Paweł Rutkowski

Between 2010 and 2020, the research team of the Section for Sign Linguistics collected, annotated, and translated a large corpus of Polish Sign Language (polski język migowy, PJM). After this task was finished, a substan…

Towards Mapping Thesauri onto plWordNet

2018-01-01 · GWC 2018 1 · Marek Maziarz, Maciej Piasecki

plWordNet, the wordnet of Polish, has become a very comprehensive description of the Polish lexical system. This paper presents a plan of its semi-automated integration with thesauri, terminological databases and ontolog…

Keyword Extraction

Inforex --- a collaborative system for text corpora annotation and analysis

2017-09-01 · RANLP 2017 9 · Micha{\l} Marci{\'n}czuk, Marcin Oleksy, Jan Koco{\'n}

We report a first major upgrade of Inforex {---} a web-based system for qualitative and collaborative text corpora annotation and analysis. Inforex is a part of Polish CLARIN infrastructure. It is integrated with a digit…

Named Entity Recognition (NER)Word Sense Disambiguation

LanguageCrawl: A Generic Tool for Building Language Models Upon Common-Crawl

2016-05-01 · LREC 2016 5 · Szymon Roziewski, Wojciech Stokowiec

The web data contains immense amount of data, hundreds of billion words are waiting to be extracted and used for language research. In this work we introduce our tool LanguageCrawl which allows NLP researchers to easily …

Language ModelingLanguage Modelling

BAN-PL: a Novel Polish Dataset of Banned Harmful and Offensive Content from Wykop.pl web service

2023-08-21 · Anna Kołos, Inez Okulska, Kinga Głąbińska, Agnieszka Karlińska 외

Since the Internet is flooded with hate, it is one of the main tasks for NLP experts to master automated online content moderation. However, advancements in this field require improved access to publicly available accura…

Specificity