Chinese-Japanese Unsupervised Neural Machine Translation Using Sub-character Level Information
Unsupervised neural machine translation (UNMT) requires only monolingual data of similar language pairs during training and can produce bi-directional translation models with relatively good performance on alphabetic languages (Lample et al., 2018). However, no research has been done to logographic language pairs. This study focuses on Chinese-Japanese UNMT trained by data containing sub-character (ideograph or stroke) level information which is decomposed from character level data. BLEU scores of both character and sub-character level systems were compared against each other and the results showed that despite the effectiveness of UNMT on character level data, sub-character level data could further enhance the performance, in which the stroke level system outperformed the ideograph level system.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationTranslationSimilar Papers 제목 키워드 기반
Character Mapping and Ad-hoc Adaptation: Edinburgh's IWSLT 2020 Open Domain Translation System
This paper describes the University of Edinburgh{'}s neural machine translation systems submitted to the IWSLT 2020 open domain Japanese$\leftrightarrow$Chinese translation task. On top of commonplace techniques like tok…
Machine TranslationTranslationImproving Character-level Japanese-Chinese Neural Machine Translation with Radicals as an Additional Input Feature
In recent years, Neural Machine Translation (NMT) has been proven to get impressive results. While some additional linguistic features of input words improve word-level NMT, any additional character features have not bee…
Machine TranslationNMTTranslationChinese Characters Mapping Table of Japanese, Traditional Chinese and Simplified Chinese
Chinese characters are used both in Japanese and Chinese, which are called Kanji and Hanzi respectively. Chinese characters contain significant semantic information, a mapping table between Kanji and Hanzi can be very us…
Cross-Lingual Information RetrievalInformation RetrievalMachine TranslationRelation+2Improving Patent Translation using Bilingual Term Extraction and Re-tokenization for Chinese--Japanese
Unlike European languages, many Asian languages like Chinese and Japanese do not have typographic boundaries in written system. Word segmentation (tokenization) that break sentences down into individual words (tokens) is…
Chinese Word SegmentationMachine TranslationSegmentationTerm Extraction+1Meta Ensemble for Japanese-Chinese Neural Machine Translation: Kyoto-U+ECNU Participation to WAT 2020
This paper describes the Japanese-Chinese Neural Machine Translation (NMT) system submitted by the joint team of Kyoto University and East China Normal University (Kyoto-U+ECNU) to WAT 2020 (Nakazawa et al.,2020). We par…
DenoisingMachine TranslationNMTTranslation