DiaBLa: A Corpus of Bilingual Spontaneous Written Dialogues for Machine Translation
We present a new English-French test set for the evaluation of Machine Translation (MT) for informal, written bilingual dialogue. The test set contains 144 spontaneous dialogues (5,700+ sentences) between native English and French speakers, mediated by one of two neural MT systems in a range of role-play settings. The dialogues are accompanied by fine-grained sentence-level judgments of MT quality, produced by the dialogue participants themselves, as well as by manually normalised versions and reference translations produced a posteriori. The motivation for the corpus is two-fold: to provide (i) a unique resource for evaluating MT models, and (ii) a corpus for the analysis of MT-mediated communication. We provide a preliminary analysis of the corpus to confirm that the participants' judgments reveal perceptible differences in MT quality between the two MT systems used.
Code (2)
Tasks
Machine TranslationSentenceTranslationSimilar Papers 제목 키워드 기반
Turn Segmentation into Utterances for Arabic Spontaneous Dialogues and Instance Messages
Text segmentation task is an essential processing task for many of Natural Language Processing (NLP) such as text summarization, text translation, dialogue language understanding, among others. Turns segmentation conside…
Dialogue UnderstandingSegmentationText SegmentationText Summarization+1PentoRef: A Corpus of Spoken References in Task-oriented Dialogues
PentoRef is a corpus of task-oriented dialogues collected in systematically manipulated settings. The corpus is multilingual, with English and German sections, and overall comprises more than 20000 utterances. The dialog…
ding-01 :ARG0: An AMR Corpus for Spontaneous French Dialogue
We present our work to build a French semantic corpus by annotating French dialogue in Abstract Meaning Representation (AMR). Specifically, we annotate the DinG corpus, consisting of transcripts of spontaneous French dia…
CoRuSS - a New Prosodically Annotated Corpus of Russian Spontaneous Speech
This paper describes speech data recording, processing and annotation of a new speech corpus CoRuSS (Corpus of Russian Spontaneous Speech), which is based on connected communicative speech recorded from 60 native Russian…
TuGeBiC: A Turkish German Bilingual Code-Switching Corpus
In this paper we describe the process of collection, transcription, and annotation of recordings of spontaneous speech samples from Turkish-German bilinguals, and the compilation of a corpus called TuGeBiC. Participants …
Language Identification