paper-with-me

홈 › Papers

Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages

2026-07-02 · A. Seza Doğruöz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li, David Ifeoluwa Adelani arxiv

LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional metrics and high correlations with human judgment, albeit mostly in English. There are now attempts to extend LLM-as-a-Judge to multilingual settings including low-resource languages. However, LLMs have limited proficiency in low-resource languages, and there is often no adequate human validation in these settings. To highlight the scope of the problem and current practices, we explore the use of LLM-as-a-Judge evaluators in ACL Anthology papers focusing on multilingual settings and low-resource languages across a diverse set of tasks. Out of 650 papers mentioning LLM-as-a-judge, only 33 of them focus on low-resource or multilingual settings. Our in-depth analysis of these papers indicates inconsistent evaluation outcomes, a tendency to overtrust LLM judgments in multilingual settings, and the widespread reliance on a single judge model per study. To help the NLP community further, we conclude with recommendations about how to use LLM-as-a-Judge in multilingual and low-resource settings.

📄 PDF Abstract BibTeX arXiv:2607.02235

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Reliable Multilingual LLMs-as-a-Judge: An Empirical Study

2026-05-27 · Irune Zubiaga, Aitor Soroa, Rodrigo Agerri arxiv

Large language models (LLMs) are increasingly used for the automatic evaluation of generated text, yet most prior work focuses on English. Despite the growing demand for multilingual evaluation, extending LLM-based evalu…

Checklist Engineering Empowers Multilingual LLM Judges

2025-07-09 · Mohammad Ghiasvand Mohammadkhani, Hamid Beigy arxiv

Automated text evaluation has long been a central issue in Natural Language Processing (NLP). Recently, the field has shifted toward using Large Language Models (LLMs) as evaluators-a trend known as the LLM-as-a-Judge pa…

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models?

2025-02-10 · Gonçalo Gomes, Chrysoula Zerva, Bruno Martins

The evaluation of image captions, looking at both linguistic fluency and semantic correspondence to visual contents, has witnessed a significant effort. Still, despite advancements such as the CLIPScore metric, multiling…

Image CaptioningSemantic correspondence

Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation

2026-04-06 · Firoj Alam, Gagan Bhatia, Sahinur Rahman Laskar, Shammur Absar Chowdhury arxiv

While Large Language Models (LLMs) are increasingly adopted as automated judges for evaluating generated text, their outputs are often costly, and highly sensitive to prompt design, language, and aggregation strategies, …

Question Answering

Code Mixologist : A Practitioner's Guide to Building Code-Mixed LLMs

2026-01-21 · Himanshu Gupta, Pratik Jayarao, Chaitanya Dwivedi, Neeraj Varshney arxiv

Code-mixing and code-switching (CSW) remain challenging phenomena for large language models (LLMs). Despite recent advances in multilingual modeling, LLMs often struggle in mixed-language settings, exhibiting systematic …