A Corpus of Native, Non-native and Translated Texts
We describe a monolingual English corpus of original and (human) translated texts, with an accurate annotation of speaker properties, including the original language of the utterances and the speaker{'}s country of origin. We thus obtain three sub-corpora of texts reflecting native English, non-native English, and English translated from a variety of European languages. This dataset will facilitate the investigation of similarities and differences between these kinds of sub-languages. Moreover, it will facilitate a unified comparative study of translations and language produced by (highly fluent) non-native speakers, two closely-related phenomena that have only been studied in isolation so far.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Building The First English-Brazilian Portuguese Corpus for Automatic Post-Editing
This paper introduces the first corpus for Automatic Post-Editing of English and a low-resource language, Brazilian Portuguese. The source English texts were extracted from the WebNLG corpus and automatically translated …
Automatic Post-EditingMachine TranslationTranslationOn the Similarities Between Native, Non-native and Translated Texts
We present a computational analysis of three language varieties: native, advanced non-native, and translation. Our goal is to investigate the similarities and differences between non-native language productions and trans…
TranslationLow-resource Information Extraction with the European Clinical Case Corpus
We present E3C-3.0, a multilingual dataset in the medical domain, comprising clinical cases annotated with diseases and test-result relations. The dataset includes both native texts in five languages (English, French, It…
Transfer LearningDialects of Translationese Shape Language Model Learning
Machine-translated data is widely used in multilingual NLP, particularly where native text is scarce. However, translated text differs systematically from native text. This phenomenon is known as translationese, and it r…
Linguistic AcceptabilityLanguage ModellingSimple TICO-19: A Dataset for Joint Translation and Simplification of COVID-19 Texts
Specialist high-quality information is typically first available in English, and it is written in a language that may be difficult to understand by most readers. While Machine Translation technologies contribute to mitig…
Machine TranslationTranslation