The Polish Sejm Corpus
This document presents the first edition of the Polish Sejm Corpus -- a new specialized resource containing transcribed, automatically annotated utterances of the Members of Polish Sejm (lower chamber of the Polish Parliament). The corpus data encoding is inherited from the National Corpus of Polish and enhanced with session metadata and structure. The multi-layered stand-off annotation contains sentence- and token-level segmentation, disambiguated morphosyntactic information, syntactic words and groups resulting from shallow parsing and named entities. The paper also outlines several novel ideas for corpus preparation, e.g. the notion of a live corpus, constantly populated with new data or the concept of linking corpus data with external databases to enrich content. Although initial statistical comparison of the resource with the balanced corpus of general Polish reveals substantial differences in language richness, the resource makes a valuable source of linguistic information as a large (300 M segments) collection of quasi-spoken data ready to be aligned with the audio/video recording of sessions, currently being made publicly available by Sejm.
Code (0)
등록된 구현이 없습니다.
Tasks
SentenceWord Sense DisambiguationSimilar Papers 제목 키워드 기반
Expanding Abbreviations in a Strongly Inflected Language: Are Morphosyntactic Tags Sufficient?
In this paper, the problem of recovery of morphological information lost in abbreviated forms is addressed with a focus on highly inflected languages. Evidence is presented that the correct inflected form of an expanded …
TAGDiaBiz.Kom - towards a Polish Dialogue Act Corpus Based on ISO 24617-2 Standard
This article presents the specification and evaluation of DiaBiz.Kom – the corpus of dialogue texts in Polish. The corpus contains transcriptions of telephone conversations conducted according to a prepared scenario. The…
The Polish Summaries Corpus
This article presents the Polish Summaries Corpus, a new resource created to support the development and evaluation of the tools for automated single-document summarization of Polish. The Corpus contains a large number o…
ArticlesDocument SummarizationOpen Repository of the Polish Sign Language Corpus: Publication Project of the Polish Sign Language Corpus
Between 2010 and 2020, the research team of the Section for Sign Linguistics collected, annotated, and translated a large corpus of Polish Sign Language (polski język migowy, PJM). After this task was finished, a substan…
New Developments in the Polish Parliamentary Corpus
This short paper presents the current (as of February 2020) state of preparation of the Polish Parliamentary Corpus (PPC){---}an extensive collection of transcripts of Polish parliamentary proceedings dating from 1919 to…