paper-with-me

Papers

WMT20 Document-Level Markable Error Exploration

2020-11-01 · WMT (EMNLP) 2020 11 · Vilém Zouhar, Tereza Vojtěchová, Ondřej Bojar

Even though sentence-centric metrics are used widely in machine translation evaluation, document-level performance is at least equally important for professional usage. In this paper, we bring attention to detailed document-level evaluation focused on markables (expressions bearing most of the document meaning) and the negative impact of various markable error phenomena on the translation. For an annotation experiment of two phases, we chose Czech and English documents translated by systems submitted to WMT20 News Translation Task. These documents are from the News, Audit and Lease domains. We show that the quality and also the kind of errors varies significantly among the domains. This systematic variance is in contrast to the automatic evaluation results. We inspect which specific markables are problematic for MT systems and conclude with an analysis of the effect of markable error types on the MT performance measured by humans and automatic evaluation tools.

📄 PDF Abstract BibTeX

Code (1)

ELITR/wmt20-elitr-testsuite 공식 구현

Tasks

Machine TranslationSentenceTranslation

Similar Papers 제목 키워드 기반

Revise: A Framework for Revising OCRed text in Practical Information Systems with Data Contamination Strategy

2026-04-09 · Gyuho Shim, Seongtae Hong, Heuiseok Lim arxiv

Recent advances in Large Language Models (LLMs) have significantly improved the field of Document AI, demonstrating remarkable performance on document understanding tasks such as question answering. However, existing app…

Synthetic Data GenerationQuestion AnsweringDocument AI

Graphy'our Data: Towards End-to-End Modeling, Exploring and Generating Report from Raw Data

2025-02-24 · Longbin Lai, Changwei Luo, Yunkai Lou, Mingchen Ju 외

Large Language Models (LLMs) have recently demonstrated remarkable performance in tasks such as Retrieval-Augmented Generation (RAG) and autonomous AI agent workflows. Yet, when faced with large sets of unstructured docu…

AI AgentRAGRetrieval-augmented GenerationSurvey

Is ChatGPT a Highly Fluent Grammatical Error Correction System? A Comprehensive Evaluation

2023-04-04 · Tao Fang, Shu Yang, Kaixin Lan, Derek F. Wong 외

ChatGPT, a large-scale language model based on the advanced GPT-3.5 architecture, has shown remarkable potential in various Natural Language Processing (NLP) tasks. However, there is currently a dearth of comprehensive s…

Grammatical Error CorrectionIn-Context LearningLanguage ModelingLanguage Modelling+1

Document-level grammatical error correction

2021-04-01 · EACL (BEA) 2021 4 · Zheng Yuan, Christopher Bryant

Document-level context can provide valuable information in grammatical error correction (GEC), which is crucial for correcting certain errors and resolving inconsistencies. In this paper, we investigate context-aware app…

Grammatical Error CorrectionNMTSentence

Text Network Exploration via Heterogeneous Web of Topics

2016-10-02 · Junxian He, Ying Huang, Changfeng Liu, Jiaming Shen 외

A text network refers to a data type that each vertex is associated with a text document and the relationship between documents is represented by edges. The proliferation of text networks such as hyperlinked webpages and…