paper-with-me

Papers

A Corpus of Spontaneous Multi-party Conversation in Bosnian Serbo-Croatian and British English

2012-05-01 · LREC 2012 5 · Emina Kurti{\'c}, Bill Wells, Guy J. Brown, Timothy Kempton, Ahmet Aker

In this paper we present a corpus of audio and video recordings of spontaneous, face-to-face multi-party conversation in two languages. Freely available high quality recordings of mundane, non-institutional, multi-party talk are still sparse, and this corpus aims to contribute valuable data suitable for study of multiple aspects of spoken interaction. In particular, it constitutes a unique resource for spoken Bosnian Serbo-Croatian (BSC), an under-resourced language with no spoken resources available at present. The corpus consists of just over 3 hours of free conversation in each of the target languages, BSC and British English (BE). The audio recordings have been made on separate channels using head-set microphones, as well as using a microphone array, containing 8 omni-directional microphones. The data has been segmented and transcribed using segmentation notions and transcription conventions developed from those of the conversation analysis research tradition. Furthermore, the transcriptions have been automatically aligned with the audio at the word and phone level, using the method of forced alignment. In this paper we describe the procedures behind the corpus creation and present the main features of the corpus for the study of conversation.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RoomReader: A Multimodal Corpus of Online Multiparty Conversational Interactions

2022-06-01 · LREC 2022 6 · Justine Reverdy, Sam O’Connor Russell, Louise Duquenne, Diego Garaialde 외

We present RoomReader, a corpus of multimodal, multiparty conversational interactions in which participants followed a collaborative student-tutor scenario designed to elicit spontaneous speech. The corpus was developed …

multimodal interaction

Discourse Structure and Dialogue Acts in Multiparty Dialogue: the STAC Corpus

2016-05-01 · LREC 2016 5 · Nicholas Asher, Julie Hunter, Mathieu Morey, Benamara Farah 외

This paper describes the STAC resource, a corpus of multi-party chats annotated for discourse structure in the style of SDRT (Asher and Lascarides, 2003; Lascarides and Asher, 2009). The main goal of the STAC project is …

UgChDial: A Uyghur Chat-based Dialogue Corpus for Response Space Classification

2022-06-01 · LREC 2022 6 · Zulipiye Yusupujiang, Jonathan Ginzburg

In this paper, we introduce a carefully designed and collected language resource: UgChDial – a Uyghur dialogue corpus based on a chatroom environment. The Uyghur Chat-based Dialogue Corpus (UgChDial) is divided into two …

ASCEND: A Spontaneous Chinese-English Dataset for Code-switching in Multi-turn Conversation

2021-12-12 · LREC 2022 6 · Holy Lovenia, Samuel Cahyawijaya, Genta Indra Winata, Peng Xu 외

Code-switching is a speech phenomenon occurring when a speaker switches language during a conversation. Despite the spontaneous nature of code-switching in conversational spoken language, most existing works collect code…

Who Is Speaking to Whom? Learning to Identify Utterance Addressee in Multi-Party Conversations

2019-11-01 · IJCNLP 2019 11 · Ran Le, Wenpeng Hu, Mingyue Shang, Zhenjun You 외

Previous research on dialogue systems generally focuses on the conversation between two participants, yet multi-party conversations which involve more than two participants within one session bring up a more complicated …