PETCI: A Parallel English Translation Dataset of Chinese Idioms
Idioms are an important language phenomenon in Chinese, but idiom translation is notoriously hard. Current machine translation models perform poorly on idiom translation, while idioms are sparse in many translation datasets. We present PETCI, a parallel English translation dataset of Chinese idioms, aiming to improve idiom translation by both human and machine. The dataset is built by leveraging human and machine effort. Baseline generation models show unsatisfactory abilities to improve translation, but structure-aware classification models show good performance on distinguishing good translations. Furthermore, the size of PETCI can be easily increased without expertise. Overall, PETCI can be helpful to language learners and machine translation systems.
Code (1)
Tasks
Machine TranslationTranslationSimilar Papers 제목 키워드 기반
Bidirectional Chinese and English Passive Sentences Dataset for Machine Translation
Machine Translation (MT) evaluation has gone beyond metrics, towards more specific linguistic phenomena. Regarding English-Chinese language pairs, passive sentences are constructed and distributed differently due to lang…
Machine TranslationNEJM-enzh: A Parallel Corpus for English-Chinese Translation in the Biomedical Domain
Machine translation requires large amounts of parallel text. While such datasets are abundant in domains such as newswire, they are less accessible in the biomedical domain. Chinese and English are two of the most widely…
Machine TranslationSentenceTranslationNICT's Supervised Neural Machine Translation Systems for the WMT19 News Translation Task
In this paper, we describe our supervised neural machine translation (NMT) systems that we developed for the news translation task for Kazakh↔English, Gujarati↔English, Chinese↔English, and English→Finnish translation di…
Machine TranslationNMTTransfer LearningTranslationAdaptive Few-shot Prompting for Machine Translation with Pre-trained Language Models
Recently, Large language models (LLMs) with in-context learning have demonstrated remarkable potential in handling neural machine translation. However, existing evidence shows that LLMs are prompt-sensitive and it is sub…
In-Context LearningMachine TranslationRetrievalTranslationThe University of Maryland's Chinese-English Neural Machine Translation Systems at WMT18
This paper describes the University of Maryland{'}s submission to the WMT 2018 Chinese↔English news translation tasks. Our systems are BPE-based self-attentional Transformer networks with parallel and backtranslated mono…
Machine TranslationRerankingTranslation