paper-with-me

홈 › Papers

CJaFr-v3 : A Freely Available Filtered Japanese-French Aligned Corpus

2022-08-28 · Raoul Blin, Fabien Cromières

We present a free Japanese-French parallel corpus. It includes 15M aligned segments and is obtained by compiling and filtering several existing resources. In this paper, we describe the existing resources, their quantity and quality, the filtering we applied to improve the quality of the corpus, and the content of the ready-to-use corpus. We also evaluate the usefulness of this corpus and the quality of our filtering by training and evaluating some standard MT systems with it.

📄 PDF Abstract BibTeX arXiv:2208.13170

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TCOF-POS : un corpus libre de fran\ccais parl\'e annot\'e en morphosyntaxe (TCOF-POS : A Freely Available POS-Tagged Corpus of Spoken French) [in French]

2012-06-01 · JEPTALNRECITAL 2012 6 · Christophe Benzitoun, Kar{\"e}n Fort, Beno{\^\i}t Sagot
POS

ANCOR, the first large French speaking corpus of conversational speech annotated in coreference to be freely available (ANCOR, premier corpus de fran\ccais parl\'e d'envergure annot\'e en cor\'ef\'erence et distribu\'e librement) [in French]

2013-06-01 · JEPTALNRECITAL 2013 6 · Judith Muzerelle, Ana{\"\i}s Lefeuvre, Jean-Yves Antoine, Emmanuel Schang 외

Pre-training via Leveraging Assisting Languages for Neural Machine Translation

2020-07-01 · ACL 2020 6 · Haiyue Song, Raj Dabre, Zhuoyuan Mao, Fei Cheng 외

Sequence-to-sequence (S2S) pre-training using large monolingual data is known to improve performance for various S2S NLP tasks. However, large monolingual corpora might not always be available for the languages of intere…

Machine TranslationNMTTranslation

Towards an Automatic Classification of Illustrative Examples in a Large Japanese-French Dictionary Obtained by OCR

2018-08-01 · COLING 2018 8 · Christian Boitet, Mathieu Mangeot, Mutsuko Tomokiyo

We work on improving the Cesselin, a large and open source Japanese-French bilingual dictionary digitalized by OCR, available on the web, and contributively improvable online. Labelling its examples (about 226000) would …

General ClassificationMachine TranslationOptical Character Recognition (OCR)

JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis

2017-10-28 · Ryosuke Sonobe, Shinnosuke Takamichi, Hiroshi Saruwatari

Thanks to improvements in machine learning techniques including deep learning, a free large-scale speech corpus that can be shared between academic institutions and commercial companies has an important role. However, su…

BIG-bench Machine LearningSpeech Synthesis