paper-with-me

홈 › Papers

The Polish Summaries Corpus

2014-05-01 · LREC 2014 5 · Maciej Ogrodniczuk, Mateusz Kope{\'c}

This article presents the Polish Summaries Corpus, a new resource created to support the development and evaluation of the tools for automated single-document summarization of Polish. The Corpus contains a large number of manual summaries of news articles, with many independently created summaries for a single text. Such approach is supposed to overcome the annotator bias, which is often described as a problem during the evaluation of the summarization algorithms against a single gold standard. There are several summarizers developed specifically for Polish language, but their in-depth evaluation and comparison was impossible without a large, manually created corpus. We present in detail the process of text selection, annotation process and the contents of the corpus, which includes both abstract free-word summaries, as well as extraction-based summaries created by selecting text spans from the original document. Finally, we describe how that resource could be used not only for the evaluation of the existing summarization tools, but also for studies on the human summarization process in Polish language.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesDocument Summarization

Similar Papers 제목 키워드 기반

The Polish Sejm Corpus

2012-05-01 · LREC 2012 5 · Maciej Ogrodniczuk

This document presents the first edition of the Polish Sejm Corpus -- a new specialized resource containing transcribed, automatically annotated utterances of the Members of Polish Sejm (lower chamber of the Polish Parli…

SentenceWord Sense Disambiguation

DiaBiz.Kom - towards a Polish Dialogue Act Corpus Based on ISO 24617-2 Standard

2022-10-01 · COLING 2022 10 · Marcin Oleksy, Jan Wieczorek, Dorota Drużyłowska, Julia Klyus 외

This article presents the specification and evaluation of DiaBiz.Kom – the corpus of dialogue texts in Polish. The corpus contains transcriptions of telephone conversations conducted according to a prepared scenario. The…

Open Repository of the Polish Sign Language Corpus: Publication Project of the Polish Sign Language Corpus

2022-06-01 · SignLang (LREC) 2022 6 · Anna Kuder, Joanna Wójcicka, Piotr Mostowski, Paweł Rutkowski

Between 2010 and 2020, the research team of the Section for Sign Linguistics collected, annotated, and translated a large corpus of Polish Sign Language (polski język migowy, PJM). After this task was finished, a substan…

New Developments in the Polish Parliamentary Corpus

2020-05-01 · LREC 2020 5 · Maciej Ogrodniczuk, Bart{\l}omiej Nito{\'n}

This short paper presents the current (as of February 2020) state of preparation of the Polish Parliamentary Corpus (PPC){---}an extensive collection of transcripts of Polish parliamentary proceedings dating from 1919 to…

HerBERT Based Language Model Detects Quantifiers and Their Semantic Properties in Polish

2022-06-01 · LREC 2022 6 · Marcin Woliński, Bartłomiej Nitoń, Witold Kieraś, Jakub Szymanik

The paper presents a tool for automatic marking up of quantifying expressions, their semantic features, and scopes. We explore the idea of using a BERT based neural model for the task (in this case HerBERT, a model train…

Language ModelingLanguage Modelling