paper-with-me

홈 › Papers

XLGoBench: Detecting cross-lingual skill gaps with algorithmic tasks

2026-05-29 · Purvam Jain, Preethi Jyothi, Vihari Piratla, Suvrat Raju arxiv

We introduce a set of synthetic algorithmic tasks to detect cross-lingual gaps in the abilities of large language models. Our benchmark is commensurate across languages, since it requires models to perform the same underlying task in different languages; scalable, since each task can be generated at varying levels of complexity allowing it to be adapted to models with different capabilities; quantifiable, since every task admits an objective notion of correctness; and transparent, since tasks are generated from simple templates that can be readily audited for translation errors. Because our benchmark focuses on algorithmic tasks, differential performance is a sufficient -- but not necessary -- indicator of cross-lingual gaps. Nevertheless, we show through extensive experiments that our benchmark exposes persistent cross-lingual gaps in multiple state-of-the-art models.

📄 PDF Abstract BibTeX arXiv:2605.30788

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BLUFF: Benchmarking the Detection of False and Synthetic Content across 58 Low-Resource Languages

2026-02-28 · Jason Lucas, Matt Murtagh-White, Adaku Uchendu, Ali Al-Lawati 외 arxiv

Multilingual falsehoods threaten information integrity worldwide, yet detection benchmarks remain confined to English or a few high-resource languages, leaving low-resource linguistic communities without robust defense t…

Mitigating the Linguistic Gap with Phonemic Representations for Robust Cross-lingual Transfer

2024-02-22 · Haeji Jung, Changdae Oh, Jooeon Kang, Jimin Sohn 외

Approaches to improving multilingual language understanding often struggle with significant performance gaps between high-resource and low-resource languages. While there are efforts to align the languages in a single la…

Cross-Lingual TransferLanguage Modelling

Cross-Language Bias Examination in Large Language Models

2025-12-17 · Yuxuan Liang, Marwa Mahmoud arxiv

This study introduces an innovative multilingual bias evaluation framework for assessing bias in Large Language Models, combining explicit bias assessment through the BBQ benchmark with implicit bias measurement using a …

Teaching LLMs to Abstain across Languages via Multilingual Feedback

2024-06-22 · Shangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding 외

Multilingual LLMs often have knowledge disparities across languages, with larger gaps in under-resourced languages. Teaching LLMs to abstain in the face of knowledge gaps is thus a promising strategy to mitigate hallucin…

Language ModelingLanguage Modelling

Cross-Lingual Response Consistency in Large Language Models: An ILR-Informed Evaluation of Claude Across Six Languages

2026-04-29 · Camelia Baluta arxiv

This paper introduces a systematic evaluation framework grounded in the Interagency Language Roundtable (ILR) Skill Level Descriptions and applies it to Claude (Sonnet 4.6) across six languages: English, French, Romanian…