paper-with-me

홈 › Papers

DiaBLa: A Corpus of Bilingual Spontaneous Written Dialogues for Machine Translation

2019-05-30 · Rachel Bawden, Sophie Rosset, Thomas Lavergne, Eric Bilinski

We present a new English-French test set for the evaluation of Machine Translation (MT) for informal, written bilingual dialogue. The test set contains 144 spontaneous dialogues (5,700+ sentences) between native English and French speakers, mediated by one of two neural MT systems in a range of role-play settings. The dialogues are accompanied by fine-grained sentence-level judgments of MT quality, produced by the dialogue participants themselves, as well as by manually normalised versions and reference translations produced a posteriori. The motivation for the corpus is two-fold: to provide (i) a unique resource for evaluating MT models, and (ii) a corpus for the analysis of MT-mediated communication. We provide a preliminary analysis of the corpus to confirm that the participants' judgments reveal perceptible differences in MT quality between the two MT systems used.

📄 PDF Abstract BibTeX arXiv:1905.13354

Code (2)

rbawden/DiaBLa-dataset 공식 구현
rbawden/diabla-chat-interface

Tasks

Machine TranslationSentenceTranslation

Similar Papers 제목 키워드 기반

Turn Segmentation into Utterances for Arabic Spontaneous Dialogues and Instance Messages

2015-05-12 · AbdelRahim A. Elmadany, Sherif M. Abdou, Mervat Gheith

Text segmentation task is an essential processing task for many of Natural Language Processing (NLP) such as text summarization, text translation, dialogue language understanding, among others. Turns segmentation conside…

Dialogue UnderstandingSegmentationText SegmentationText Summarization+1

PentoRef: A Corpus of Spoken References in Task-oriented Dialogues

2016-05-01 · LREC 2016 5 · Sina Zarrie{\ss}, Julian Hough, Casey Kennington, Ramesh Manuvinakurike 외

PentoRef is a corpus of task-oriented dialogues collected in systematically manipulated settings. The corpus is multilingual, with English and German sections, and overall comprises more than 20000 utterances. The dialog…

ding-01 :ARG0: An AMR Corpus for Spontaneous French Dialogue

2025-08-18 · Jeongwoo Kang, Maria Boritchev, Maximin Coavoux arxiv

We present our work to build a French semantic corpus by annotating French dialogue in Abstract Meaning Representation (AMR). Specifically, we annotate the DinG corpus, consisting of transcripts of spontaneous French dia…

CoRuSS - a New Prosodically Annotated Corpus of Russian Spontaneous Speech

2016-05-01 · LREC 2016 5 · Tatiana Kachkovskaia, Daniil Kocharov, Pavel Skrelin, Nina Volskaya

This paper describes speech data recording, processing and annotation of a new speech corpus CoRuSS (Corpus of Russian Spontaneous Speech), which is based on connected communicative speech recorded from 60 native Russian…

TuGeBiC: A Turkish German Bilingual Code-Switching Corpus

2022-05-02 · Jeanine Treffers-Daller and, Ozlem Çetinoğlu

In this paper we describe the process of collection, transcription, and annotation of recordings of spontaneous speech samples from Turkish-German bilinguals, and the compilation of a corpus called TuGeBiC. Participants …

Language Identification