paper-with-me

홈 › Papers

Test Set Quality in Multilingual LLM Evaluation

2025-08-04 · Chalamalasetti Kranti, Gabriel Bernier-Colborne, Yvan Gauthier, Sowmya Vajjala arxiv

Several multilingual benchmark datasets have been developed in a semi-automatic manner in the recent past to measure progress and understand the state-of-the-art in the multilingual capabilities of Large Language Models. However, there is not a lot of attention paid to the quality of the datasets themselves, despite the existence of previous work in identifying errors in even fully human-annotated test sets. In this paper, we manually analyze recent multilingual evaluation sets in two languages - French and Telugu, identifying several errors in the process. We compare the performance difference across several LLMs with the original and revised versions of the datasets and identify large differences (almost 10% in some cases) in both languages). Based on these results, we argue that test sets should not be considered immutable and should be revisited, checked for correctness, and potentially versioned. We end with some recommendations for both the dataset creators as well as consumers on addressing the dataset quality issues.

📄 PDF Abstract BibTeX arXiv:2508.02635

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models?

2025-02-10 · Gonçalo Gomes, Chrysoula Zerva, Bruno Martins

The evaluation of image captions, looking at both linguistic fluency and semantic correspondence to visual contents, has witnessed a significant effort. Still, despite advancements such as the CLIPScore metric, multiling…

Image CaptioningSemantic correspondence

The MuCoW Test Suite at WMT 2019: Automatically Harvested Multilingual Contrastive Word Sense Disambiguation Test Sets for Machine Translation

2019-08-01 · WS 2019 8 · Aless Raganato, ro, Yves Scherrer, J{\"o}rg Tiedemann

Supervised Neural Machine Translation (NMT) systems currently achieve impressive translation quality for many language pairs. One of the key features of a correct translation is the ability to perform word sense disambig…

Machine TranslationNMTSentenceTranslation+1

TenTrans Multilingual Low-Resource Translation System for WMT21 Indo-European Languages Task

2021-11-01 · WMT (EMNLP) 2021 11 · Han Yang, Bojie Hu, Wanying Xie, Ambyera Han 외

This paper describes TenTrans’ submission to WMT21 Multilingual Low-Resource Translation shared task for the Romance language pairs. This task focuses on improving translation quality from Catalan to Occitan, Romanian an…

Transfer LearningTranslation

From Synthetic to Native: Benchmarking Multilingual Intent Classification in Logistics Customer Service

2026-03-24 · Haoyu He, Jinyu Zhuang, Haoran Chu, Shuhang Yu 외 arxiv

Multilingual intent classification is central to customer-service systems on global logistics platforms, where models must process noisy user queries across languages and hierarchical label spaces. Yet most existing mult…

Cross-Lingual TransferIntent Classification

Model and Evaluation: Towards Fairness in Multilingual Text Classification

2023-03-28 · Nankai Lin, Junheng He, Zhenghang Tang, Dong Zhou 외

Recently, more and more research has focused on addressing bias in text classification models. However, existing research mainly focuses on the fairness of monolingual text classification models, and research on fairness…

ClassificationContrastive LearningFairnessLanguage Modelling+3