paper-with-me

Papers

Automatic Evaluation Metrics for Document-level Translation: Overview, Challenges and Trends

2025-04-21 · Jiaxin Guo, Xiaoyu Chen, Zhiqiang Rao, Jinlong Yang, Zongyao Li, Hengchao Shang, Daimeng Wei, Hao Yang

With the rapid development of deep learning technologies, the field of machine translation has witnessed significant progress, especially with the advent of large language models (LLMs) that have greatly propelled the advancement of document-level translation. However, accurately evaluating the quality of document-level translation remains an urgent issue. This paper first introduces the development status of document-level translation and the importance of evaluation, highlighting the crucial role of automatic evaluation metrics in reflecting translation quality and guiding the improvement of translation systems. It then provides a detailed analysis of the current state of automatic evaluation schemes and metrics, including evaluation methods with and without reference texts, as well as traditional metrics, Model-based metrics and LLM-based metrics. Subsequently, the paper explores the challenges faced by current evaluation methods, such as the lack of reference diversity, dependence on sentence-level alignment information, and the bias, inaccuracy, and lack of interpretability of the LLM-as-a-judge method. Finally, the paper looks ahead to the future trends in evaluation methods, including the development of more user-friendly document-level evaluation methods and more robust LLM-as-a-judge methods, and proposes possible research directions, such as reducing the dependency on sentence-level information, introducing multi-level and multi-granular evaluation approaches, and training models specifically for machine translation evaluation. This study aims to provide a comprehensive analysis of automatic evaluation for document-level translation and offer insights into future developments.

📄 PDF Abstract BibTeX arXiv:2504.14804

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentenceTranslation

Similar Papers 제목 키워드 기반

BlonDe: An Automatic Evaluation Metric for Document-level Machine Translation

2021-03-22 · NAACL 2022 7 · Yuchen Eleanor Jiang, Tianyu Liu, Shuming Ma, Dongdong Zhang 외

Standard automatic metrics, e.g. BLEU, are not reliable for document-level MT evaluation. They can neither distinguish document-level improvements in translation quality from sentence-level ones, nor identify the discour…

Document Level Machine TranslationMachine TranslationSentenceTranslation

{BlonDe}: An Automatic Evaluation Metric for Document-level Machine Translation

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Standard automatic metrics, e.g. BLEU, are not reliable for document-level MT evaluation. They can neither distinguish document-level improvements in translation quality from sentence-level ones, nor identify the discour…

Document Level Machine TranslationMachine TranslationSentenceTranslation

Extending Automatic Machine Translation Evaluation to Book-Length Documents

2025-09-21 · Kuang-Da Wang, Shuoyang Ding, Chao-Han Huck Yang, Ping-Chun Hsieh 외 arxiv

Despite Large Language Models (LLMs) demonstrating superior translation performance and long-context capabilities, evaluation methodologies remain constrained to sentence-level assessment due to dataset limitations, toke…

Machine Translation

DELA Project: Document-level Machine Translation Evaluation

2022-06-01 · EAMT 2022 6 · Sheila Castilho

This paper presents the results of the DELA Project. We describe the testing of context span for document-level evaluation, construction of a document-level corpus, and context position, as well as the latest development…

Document Level Machine TranslationMachine TranslationPositionTranslation

Exploring the Importance of Source Text in Automatic Post-Editing for Context-Aware Machine Translation

2021-05-01 · NoDaLiDa 2021 5 · Chaojun Wang, Christian Hardmeier, Rico Sennrich

Accurate translation requires document-level information, which is ignored by sentence-level machine translation. Recent work has demonstrated that document-level consistency can be improved with automatic post-editing (…

Automatic Post-EditingMachine TranslationSentenceTranslation