paper-with-me

홈 › Papers

Design and Evaluation of the Corpus of Everyday Japanese Conversation

2022-06-01 · LREC 2022 6 · Hanae Koiso, Haruka Amatani, Yasuharu Den, Yuriko Iseki, Yuichi Ishimoto, Wakako Kashino, Yoshiko Kawabata, Ken’ya Nishikawa, Yayoi Tanaka, Yasuyuki Usuda, Yuka Watanabe

We have constructed the Corpus of Everyday Japanese Conversation (CEJC) and published it in March 2022. The CEJC is designed to contain various kinds of everyday conversations in a balanced manner to capture their diversity. The CEJC features not only audio but also video data to facilitate precise understanding of the mechanism of real-life social behavior. The publication of a large-scale corpus of everyday conversations that includes video data is a new approach. The CEJC contains 200 hours of speech, 577 conversations, about 2.4 million words, and a total of 1675 conversants. In this paper, we present an overview of the corpus, including the recording method and devices, structure of the corpus, formats of video and audio files, transcription, and annotations. We then report some results of the evaluation of the CEJC in terms of conversant and conversation attributes. We show that the CEJC includes a good balance of adult conversants in terms of gender and age, as well as a variety of conversations in terms of conversation forms, places, activities, and numbers of conversants.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Diversity

Similar Papers 제목 키워드 기반

Survey of Conversational Behavior: Towards the Design of a Balanced Corpus of Everyday Japanese Conversation

2016-05-01 · LREC 2016 5 · Hanae Koiso, Tomoyuki Tsuchiya, Ryoko Watanabe, Daisuke Yokomori 외

In 2016, we set about building a large-scale corpus of everyday Japanese conversation―a collection of conversations embedded in naturally occurring activities in daily life. We will collect more than 200 hours of recor…

Survey

Construction of the Corpus of Everyday Japanese Conversation: An Interim Report

2018-05-01 · LREC 2018 5 · Hanae Koiso, Yasuharu Den, Yuriko Iseki, Wakako Kashino 외

Document-aligned Japanese-English Conversation Parallel Corpus

2020-12-11 · WMT (EMNLP) 2020 11 · Matīss Rikters, Ryokan Ri, Tong Li, Toshiaki Nakazawa

Sentence-level (SL) machine translation (MT) has reached acceptable quality for many high-resourced languages, but not document-level (DL) MT, which is difficult to 1) train with little amount of DL data; and 2) evaluate…

Machine TranslationSentenceTranslation

Annotated Corpora for Word Alignment between Japanese and English and its Evaluation with MAP-based Word Aligner

2012-05-01 · LREC 2012 5 · Tsuyoshi Okita

This paper presents two annotated corpora for word alignment between Japanese and English. We annotated on top of the IWSLT-2006 and the NTCIR-8 corpora. The IWSLT-2006 corpus is in the domain of travel conversation whil…

Machine TranslationSentenceWord Alignment

A Colloquial Corpus of Japanese Sign Language: Linguistic Resources for Observing Sign Language Conversations

2014-05-01 · LREC 2014 5 · Mayumi Bono, Kouhei Kikuchi, Paul Cibulka, Yutaka Osugi

We began building a corpus of Japanese Sign Language (JSL) in April 2011. The purpose of this project was to increase awareness of sign language as a distinctive language in Japan. This corpus is beneficial not only to l…