paper-with-me

홈 › Papers

Document-aligned Japanese-English Conversation Parallel Corpus

2020-12-11 · WMT (EMNLP) 2020 11 · Matīss Rikters, Ryokan Ri, Tong Li, Toshiaki Nakazawa

Sentence-level (SL) machine translation (MT) has reached acceptable quality for many high-resourced languages, but not document-level (DL) MT, which is difficult to 1) train with little amount of DL data; and 2) evaluate, as the main methods and data sets focus on SL evaluation. To address the first issue, we present a document-aligned Japanese-English conversation corpus, including balanced, high-quality business conversation data for tuning and testing. As for the second issue, we manually identify the main areas where SL MT fails to produce adequate translations in lack of context. We then create an evaluation set where these phenomena are annotated to alleviate automatic evaluation of DL systems. We train MT models using our corpus to demonstrate how using context leads to improvements.

📄 PDF Abstract BibTeX arXiv:2012.06143

Code (1)

tsuruoka-lab/AMI-Meeting-Parallel-Corpus 공식 구현

Tasks

Machine TranslationSentenceTranslation

Similar Papers 제목 키워드 기반

Context-aware Decoder for Neural Machine Translation using a Target-side Document-Level Language Model

2020-10-24 · NAACL 2021 4 · Amane Sugiyama, Naoki Yoshinaga

Although many context-aware neural machine translation models have been proposed to incorporate contexts in translation, most of those models are trained end-to-end on parallel documents aligned in sentence-level. Becaus…

DecoderLanguage ModelingLanguage ModellingMachine Translation+2

TDDC: Timely Disclosure Documents Corpus

2020-05-01 · LREC 2020 5 · Nobushige Doi, Yusuke Oda, Toshiaki Nakazawa

In this paper, we describe the details of the Timely Disclosure Documents Corpus (TDDC). TDDC was prepared by manually aligning the sentences from past Japanese and English timely disclosure documents in PDF format publi…

Machine TranslationTranslation

JESC: Japanese-English Subtitle Corpus

2017-10-29 · LREC 2018 5 · Reid Pryzant, Yongjoo Chung, Dan Jurafsky, Denny Britz

In this paper we describe the Japanese-English Subtitle Corpus (JESC). JESC is a large Japanese-English parallel corpus covering the underrepresented domain of conversational dialogue. It consists of more than 3.2 millio…

Machine TranslationTranslation

JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus

2025-08-22 · Masaaki Nagata, Katsuki Chousa, Norihito Yasuda arxiv

We constructed JaParaPat (Japanese-English Parallel Patent Application Corpus), a bilingual corpus of more than 300 million Japanese-English sentence pairs from patent applications published in Japan and the United State…

NAIST-SIC-Aligned: an Aligned English-Japanese Simultaneous Interpretation Corpus

2023-04-23 · Jinming Zhao, Yuka Ko, Kosuke Doi, Ryo Fukuda 외

It remains a question that how simultaneous interpretation (SI) data affects simultaneous machine translation (SiMT). Research has been limited due to the lack of a large-scale training corpus. In this work, we aim to fi…

Machine TranslationSentenceTranslation