paper-with-me

홈 › Papers

Open Repository of the Polish Sign Language Corpus: Publication Project of the Polish Sign Language Corpus

2022-06-01 · SignLang (LREC) 2022 6 · Anna Kuder, Joanna Wójcicka, Piotr Mostowski, Paweł Rutkowski

Between 2010 and 2020, the research team of the Section for Sign Linguistics collected, annotated, and translated a large corpus of Polish Sign Language (polski język migowy, PJM). After this task was finished, a substantial part of the gathered materials was published online as the Open Repository of the Polish Sign Language Corpus. The current paper gives an overview of the process of converting the material from the Corpus into the Repository. If presents and explains the decisions made along the way and describes the process of data preparation and publication. There are two levels of access to the Repository, which are meant to fulfil the needs of a wide range of public users, from members of the Deaf community, through hearing students of PJM, sign language teachers and interpreters, to users with academic background. We describe how corpus material available in open access was prepared to be searchable by text type and elicitation tasks, by sociolinguistic metadata, and by translation into written Polish. We go on to explain how access for research purposes differs from open access. We present possible ways in which data gathered in the Repository may be used by members of the signing community in Poland and abroad.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards a comprehensive open repository of Polish language resources

2012-05-01 · LREC 2012 5 · Maciej Ogrodniczuk, Piotr P{\k{e}}zik, Adam Przepi{\'o}rkowski

The aim of this paper is to present current efforts towards the creation of a comprehensive open repository of Polish language resources and tools (LRTs). The work described here is carried out within the CESAR project, …

LanguageCrawl: A Generic Tool for Building Language Models Upon Common-Crawl

2016-05-01 · LREC 2016 5 · Szymon Roziewski, Wojciech Stokowiec

The web data contains immense amount of data, hundreds of billion words are waiting to be extracted and used for language research. In this work we introduce our tool LanguageCrawl which allows NLP researchers to easily …

Language ModelingLanguage Modelling

PLLuM: A Family of Polish Large Language Models

2025-11-05 · Jan Kocoń, Maciej Piasecki, Arkadiusz Janz, Teddy Ferdinan 외 arxiv

Large Language Models (LLMs) play a central role in modern artificial intelligence, yet their development has been primarily focused on English, resulting in limited support for other languages. We present PLLuM (Polish …

The Polish Sejm Corpus

2012-05-01 · LREC 2012 5 · Maciej Ogrodniczuk

This document presents the first edition of the Polish Sejm Corpus -- a new specialized resource containing transcribed, automatically annotated utterances of the Members of Polish Sejm (lower chamber of the Polish Parli…

SentenceWord Sense Disambiguation

Bielik v3 Small: Technical Report

2025-05-05 · Krzysztof Ociepa, Łukasz Flis, Remigiusz Kinas, Krzysztof Wróbel 외

We introduce Bielik v3, a series of parameter-efficient generative text models (1.5B and 4.5B) optimized for Polish language processing. These models demonstrate that smaller, well-optimized architectures can achieve per…

Language ModelingLanguage Modelling