paper-with-me

홈 › Papers

DiaBiz – an Annotated Corpus of Polish Call Center Dialogs

2022-06-01 · LREC 2022 6 · Piotr Pęzik, Gosia Krawentek, Sylwia Karasińska, Paweł Wilk, Paulina Rybińska, Anna Cichosz, Angelika Peljak-Łapińska, Mikołaj Deckert, Michał Adamczyk

This paper introduces DiaBiz, a large, annotated, multimodal corpus of Polish telephone conversations conducted in varied business settings, comprising 4036 call centre interactions from nine different domains, i.e. banking, energy services, telecommunications, insurance, medical care, debt collection, tourism, retail and car rental. The corpus was developed to boost the development of third-party speech recognition engines, dialog systems and conversational intelligence tools for Polish. Its current size amounts to nearly 410 hours of recordings and over 3 million words of transcribed speech. We present the structure of the corpus, data collection and transcription procedures, challenges of punctuating and truecasing speech transcripts, dialog structure annotation and discuss some of the ecological validity considerations involved in the development of such resources.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

DiaBiz.Kom - towards a Polish Dialogue Act Corpus Based on ISO 24617-2 Standard

2022-10-01 · COLING 2022 10 · Marcin Oleksy, Jan Wieczorek, Dorota Drużyłowska, Julia Klyus 외

This article presents the specification and evaluation of DiaBiz.Kom – the corpus of dialogue texts in Polish. The corpus contains transcriptions of telephone conversations conducted according to a prepared scenario. The…

Annotation of metaphorical expressions in the Basic Corpus of Polish Metaphors

2022-06-01 · LREC 2022 6 · Elżbieta Hajnicz

This paper presents a corpus of Polish texts annotated with metaphorical expressions. It is composed of two parts of comparable size, selected from two subcorpora of the Polish National Corpus: the subcorpus manually ann…

HerBERT Based Language Model Detects Quantifiers and Their Semantic Properties in Polish

2022-06-01 · LREC 2022 6 · Marcin Woliński, Bartłomiej Nitoń, Witold Kieraś, Jakub Szymanik

The paper presents a tool for automatic marking up of quantifying expressions, their semantic features, and scopes. We explore the idea of using a BERT based neural model for the task (in this case HerBERT, a model train…

Language ModelingLanguage Modelling

The Polish Sejm Corpus

2012-05-01 · LREC 2012 5 · Maciej Ogrodniczuk

This document presents the first edition of the Polish Sejm Corpus -- a new specialized resource containing transcribed, automatically annotated utterances of the Members of Polish Sejm (lower chamber of the Polish Parli…

SentenceWord Sense Disambiguation

Polish Corpus of Annotated Descriptions of Images

2018-05-01 · LREC 2018 5 · Alina Wr{\'o}blewska
Image ClassificationImage GenerationImage RetrievalText Generation+1