paper-with-me

Papers

A Test Suite for Evaluating Discourse Phenomena in Document-level Neural Machine Translation

2020-12-01 · AACL (iwdp) 2020 12 · Xinyi Cai, Deyi Xiong

The need to evaluate the ability of context-aware neural machine translation (NMT) models in dealing with specific discourse phenomena arises in document-level NMT. However, test sets that satisfy this need are rare. In this paper, we propose a test suite to evaluate three common discourse phenomena in English-Chinese translation: pronoun, discourse connective and ellipsis where discourse divergences lie across the two languages. The test suite contains 1,200 instances, 400 for each type of discourse phenomena. We perform both automatic and human evaluation with three state-of-the-art context-aware NMT models on the proposed test suite. Results suggest that our test suite can be used as a challenging benchmark test bed for evaluating document-level NMT. The test suite will be publicly available soon.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationNMTTranslation

Similar Papers 제목 키워드 기반

Evaluating Discourse Cohesion in Pre-trained Language Models

2025-03-08 · COLING (CODI, CRAC) 2022 10 · Jie He, Wanqiu Long, Deyi Xiong

Large pre-trained neural models have achieved remarkable success in natural language process (NLP), inspiring a growing body of research analyzing their ability from different aspects. In this paper, we propose a test su…

A Test Suite and Manual Evaluation of Document-Level NMT at WMT19

2019-08-08 · Kateřina Rysová, Magdaléna Rysová, Tomáš Musil, Lucie Poláková 외

As the quality of machine translation rises and neural machine translation (NMT) is moving from sentence to document level translations, it is becoming increasingly difficult to evaluate the output of translation systems…

Machine TranslationNMTSentenceTranslation

A Test Suite and Manual Evaluation of Document-Level NMT at WMT19

2019-08-01 · WS 2019 8 · Kate{\v{r}}ina Rysov{\'a}, Magdal{\'e}na Rysov{\'a}, Tom{\'a}{\v{s}} Musil, Lucie Pol{\'a}kov{\'a} 외

As the quality of machine translation rises and neural machine translation (NMT) is moving from sentence to document level translations, it is becoming increasingly difficult to evaluate the output of translation systems…

Machine TranslationNMTSentenceTranslation

BeDiscovER: The Benchmark of Discourse Understanding in the Era of Reasoning Language Models

2025-11-17 · Chuyuan Li, Giuseppe Carenini arxiv

We introduce BeDiscovER (Benchmark of Discourse Understanding in the Era of Reasoning Language Models), an up-to-date, comprehensive suite for evaluating the discourse-level knowledge of modern LLMs. BeDiscovER compiles …

Temporal Relation ExtractionRelation ClassificationDiscourse Parsing

Disco-Bench: A Discourse-Aware Evaluation Benchmark for Language Modelling

2023-07-16 · Longyue Wang, Zefeng Du, Donghuai Liu, Deng Cai 외

Modeling discourse -- the linguistic phenomena that go beyond individual sentences, is a fundamental yet challenging aspect of natural language processing (NLP). However, existing evaluation benchmarks primarily focus on…

DiagnosticLanguage ModellingSentence