CJaFr-v3 : A Freely Available Filtered Japanese-French Aligned Corpus
We present a free Japanese-French parallel corpus. It includes 15M aligned segments and is obtained by compiling and filtering several existing resources. In this paper, we describe the existing resources, their quantity and quality, the filtering we applied to improve the quality of the corpus, and the content of the ready-to-use corpus. We also evaluate the usefulness of this corpus and the quality of our filtering by training and evaluating some standard MT systems with it.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
TCOF-POS : un corpus libre de fran\ccais parl\'e annot\'e en morphosyntaxe (TCOF-POS : A Freely Available POS-Tagged Corpus of Spoken French) [in French]
ANCOR, the first large French speaking corpus of conversational speech annotated in coreference to be freely available (ANCOR, premier corpus de fran\ccais parl\'e d'envergure annot\'e en cor\'ef\'erence et distribu\'e librement) [in French]
Pre-training via Leveraging Assisting Languages for Neural Machine Translation
Sequence-to-sequence (S2S) pre-training using large monolingual data is known to improve performance for various S2S NLP tasks. However, large monolingual corpora might not always be available for the languages of intere…
Machine TranslationNMTTranslationTowards an Automatic Classification of Illustrative Examples in a Large Japanese-French Dictionary Obtained by OCR
We work on improving the Cesselin, a large and open source Japanese-French bilingual dictionary digitalized by OCR, available on the web, and contributively improvable online. Labelling its examples (about 226000) would …
General ClassificationMachine TranslationOptical Character Recognition (OCR)JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis
Thanks to improvements in machine learning techniques including deep learning, a free large-scale speech corpus that can be shared between academic institutions and commercial companies has an important role. However, su…
BIG-bench Machine LearningSpeech Synthesis