paper-with-me

홈 › Papers

The Polish Sejm Corpus

2012-05-01 · LREC 2012 5 · Maciej Ogrodniczuk

This document presents the first edition of the Polish Sejm Corpus -- a new specialized resource containing transcribed, automatically annotated utterances of the Members of Polish Sejm (lower chamber of the Polish Parliament). The corpus data encoding is inherited from the National Corpus of Polish and enhanced with session metadata and structure. The multi-layered stand-off annotation contains sentence- and token-level segmentation, disambiguated morphosyntactic information, syntactic words and groups resulting from shallow parsing and named entities. The paper also outlines several novel ideas for corpus preparation, e.g. the notion of a live corpus, constantly populated with new data or the concept of linking corpus data with external databases to enrich content. Although initial statistical comparison of the resource with the balanced corpus of general Polish reveals substantial differences in language richness, the resource makes a valuable source of linguistic information as a large (300 M segments) collection of quasi-spoken data ready to be aligned with the audio/video recording of sessions, currently being made publicly available by Sejm.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceWord Sense Disambiguation

Similar Papers 제목 키워드 기반

Expanding Abbreviations in a Strongly Inflected Language: Are Morphosyntactic Tags Sufficient?

2017-08-20 · Piotr Żelasko

In this paper, the problem of recovery of morphological information lost in abbreviated forms is addressed with a focus on highly inflected languages. Evidence is presented that the correct inflected form of an expanded …

TAG

DiaBiz.Kom - towards a Polish Dialogue Act Corpus Based on ISO 24617-2 Standard

2022-10-01 · COLING 2022 10 · Marcin Oleksy, Jan Wieczorek, Dorota Drużyłowska, Julia Klyus 외

This article presents the specification and evaluation of DiaBiz.Kom – the corpus of dialogue texts in Polish. The corpus contains transcriptions of telephone conversations conducted according to a prepared scenario. The…

The Polish Summaries Corpus

2014-05-01 · LREC 2014 5 · Maciej Ogrodniczuk, Mateusz Kope{\'c}

This article presents the Polish Summaries Corpus, a new resource created to support the development and evaluation of the tools for automated single-document summarization of Polish. The Corpus contains a large number o…

ArticlesDocument Summarization

Open Repository of the Polish Sign Language Corpus: Publication Project of the Polish Sign Language Corpus

2022-06-01 · SignLang (LREC) 2022 6 · Anna Kuder, Joanna Wójcicka, Piotr Mostowski, Paweł Rutkowski

Between 2010 and 2020, the research team of the Section for Sign Linguistics collected, annotated, and translated a large corpus of Polish Sign Language (polski język migowy, PJM). After this task was finished, a substan…

New Developments in the Polish Parliamentary Corpus

2020-05-01 · LREC 2020 5 · Maciej Ogrodniczuk, Bart{\l}omiej Nito{\'n}

This short paper presents the current (as of February 2020) state of preparation of the Polish Parliamentary Corpus (PPC){---}an extensive collection of transcripts of Polish parliamentary proceedings dating from 1919 to…