paper-with-me

홈 › Papers

An Aligned French-Chinese corpus of 10K segments from university educational material

2016-12-01 · WS 2016 12 · Ruslan Kalitvianski, Lingxiao Wang, Val{\'e}rie Bellynck, Christian Boitet

This paper describes a corpus of nearly 10K French-Chinese aligned segments, produced by post-editing machine translated computer science courseware. This corpus was built from 2013 to 2016 within the PROJECT{\_}NAME project, by native Chinese students. The quality, as judged by native speakers, is ad-equate for understanding (far better than by reading only the original French) and for getting better marks. This corpus is annotated at segment-level by a self-assessed quality score. It has been directly used as supplemental training data to build a statistical machine translation system dedicated to that sublanguage, and can be used to extract the specific bilingual terminology. To our knowledge, it is the first corpus of this kind to be released.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

CJaFr-v3 : A Freely Available Filtered Japanese-French Aligned Corpus

2022-08-28 · Raoul Blin, Fabien Cromières

We present a free Japanese-French parallel corpus. It includes 15M aligned segments and is obtained by compiling and filtering several existing resources. In this paper, we describe the existing resources, their quantity…

European Union Language Resources in Sketch Engine

2016-05-01 · LREC 2016 5 · V{\'\i}t Baisa, Jan Michelfeit, Marek Medve{\v{d}}, Milo{\v{s}} Jakub{\'\i}{\v{c}}ek

Several parallel corpora built from European Union language resources are presented here. They were processed by state-of-the-art tools and made available for researchers in the corpus manager Sketch Engine. A completely…

The United Nations Parallel Corpus v1.0

2016-05-01 · LREC 2016 5 · Micha{\l} Ziemski, Marcin Junczys-Dowmunt, Bruno Pouliquen

This paper describes the creation process and statistics of the official United Nations Parallel Corpus, the first parallel corpus composed from United Nations documents published by the original data creator. The parall…

Translation

Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation Evaluation

2018-02-09 · LREC 2018 5 · Ali Can Kocabiyikoglu, Laurent Besacier, Olivier Kraif

Recent works in spoken language translation (SLT) have attempted to build end-to-end speech-to-text translation without using source language transcription during learning or decoding. However, while large quantities of …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationSentence+5

Using Word Embeddings to Translate Named Entities

2016-05-01 · LREC 2016 5 · Octavia-Maria {\c{S}}ulea, Sergiu Nisioi, Liviu P. Dinu

In this paper we investigate the usefulness of neural word embeddings in the process of translating Named Entities (NEs) from a resource-rich language to a language low on resources relevant to the task at hand, introduc…

Chinese Named Entity Recognitionnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+4