paper-with-me

홈 › Papers

XQ-MEval: A Dataset with Cross-lingual Parallel Quality for Benchmarking Translation Metrics

2026-04-16 · Jingxuan Liu, Zhi Qu, Jin Tei, Hidetaka Kamigaito, Lemao Liu, Taro Watanabe arxiv

Automatic evaluation metrics are essential for building multilingual translation systems. The common practice of evaluating these systems is averaging metric scores across languages, yet this is suspicious since metrics may suffer from cross-lingual scoring bias, where translations of equal quality receive different scores across languages. This problem has not been systematically studied because no benchmark exists that provides parallel-quality instances across languages, and expert annotation is not realistic. In this work, we propose XQ-MEval, a semi-automatically built dataset covering nine translation directions, to benchmark translation metrics. Specifically, we inject MQM-defined errors into gold translations automatically, filter them by native speakers for reliability, and merge errors to generate pseudo translations with controllable quality. These pseudo translations are then paired with corresponding sources and references to form triplets used in assessing the qualities of translation metrics. Using XQ-MEval, our experiments on nine representative metrics reveal the inconsistency between averaging and human judgment and provide the first empirical evidence of cross-lingual scoring bias. Finally, we propose a normalization strategy derived from XQ-MEval that aligns score distributions across languages, improving the fairness and reliability of multilingual metric evaluation.

📄 PDF Abstract BibTeX arXiv:2604.14934

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs

2024-11-14 · Yidan Zhang, Yu Wan, Boyi Deng, Baosong Yang 외

Recent advancements in large language models (LLMs) showcase varied multilingual capabilities across tasks like translation, code generation, and reasoning. Previous assessments often limited their scope to fundamental n…

Code GenerationTransfer Learning

Citius at SemEval-2017 Task 2: Cross-Lingual Similarity from Comparable Corpora and Dependency-Based Contexts

2017-08-01 · SEMEVAL 2017 8 · Pablo Gamallo

This article describes the distributional strategy submitted by the Citius team to the SemEval 2017 Task 2. Even though the team participated in two subtasks, namely monolingual and cross-lingual word similarity, the art…

Task 2Word Similarity

RUFINO at SemEval-2017 Task 2: Cross-lingual lexical similarity by extending PMI and word embeddings systems with a Swadesh's-like list

2017-08-01 · SEMEVAL 2017 8 · Sergio Jimenez, George Due{\~n}as, Lorena Gaitan, Jorge Segura

The RUFINO team proposed a non-supervised, conceptually-simple and low-cost approach for addressing the Multilingual and Cross-lingual Semantic Word Similarity challenge at SemEval 2017. The proposed systems were cross-l…

Semantic Textual SimilarityTask 2Word EmbeddingsWord Sense Disambiguation+1

UAlberta at SemEval-2020 Task 2: Using Translations to Predict Cross-Lingual Entailment

2020-12-01 · SEMEVAL 2020 · Bradley Hauer, Amir Ahmad Habibi, Yixing Luan, Arnob Mallik 외

We investigate the hypothesis that translations can be used to identify cross-lingual lexical entailment. We propose novel methods that leverage parallel corpora, word embeddings, and multilingual lexical resources. Our …

Lexical EntailmentTask 2Word Embeddings

ConceptNet at SemEval-2017 Task 2: Extending Word Embeddings with Multilingual Relational Knowledge

2017-04-11 · SEMEVAL 2017 8 · Robyn Speer, Joanna Lowry-Duda

This paper describes Luminoso's participation in SemEval 2017 Task 2, "Multilingual and Cross-lingual Semantic Word Similarity", with a system based on ConceptNet. ConceptNet is an open, multilingual knowledge graph that…

General KnowledgeMultilingual Word EmbeddingsTask 2Word Embeddings+1