paper-with-me

홈 › Papers

Idiom Understanding as a Tool to Measure the Dialect Gap

2025-10-06 · David Beauchemin, Yan Tremblay, Mohamed Amine Youssef, Richard Khoury arxiv

The tasks of idiom understanding and dialect understanding are both well-established benchmarks in natural language processing. In this paper, we propose combining them, and using regional idioms as a test of dialect understanding. Towards this end, we propose three new benchmark datasets for the Quebec dialect of French: QFrCoRE, which contains 4,633 instances of idiomatic phrases, and QFrCoRT, which comprises 171 regional instances of idiomatic words, and a new benchmark for French Metropolitan expressions, MFrCoE, which comprises 4,938 phrases. We explain how to construct these corpora, so that our methodology can be replicated for other dialects. Our experiments with 111 LLMs reveal a critical disparity in dialectal competence: while models perform well on French Metropolitan, 65.77% of them perform significantly worse on Quebec idioms, with only 9.0% favoring the regional dialect. These results confirm that our benchmarks are a reliable tool for quantifying the dialect gap and that prestige-language proficiency does not guarantee regional dialect understanding.

📄 PDF Abstract BibTeX arXiv:2510.05026

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Tools for Building a Corpus to Study the Historical and Geographical Variation of the Romanian Language

2017-09-01 · RANLP 2017 9 · Victoria Bobicev, C{\u{a}}t{\u{a}}lina M{\u{a}}r{\u{a}}nduc, Cenel Augusto Perez

Contemporary standard language corpora are ideal for NLP. There are few morphologically and syntactically annotated corpora for Romanian, and those existing or in progress only deal with the Contemporary Romanian standar…

Building a Corpus of Qatari Arabic Expressions

2020-05-01 · LREC 2020 5 · Sara Al-Mulla, Wajdi Zaghouani

The current Arabic natural language processing resources are mainly build to address the Modern Standard Arabic (MSA), while we witnessed some scattered efforts to build resources for various Arabic dialects such as the …

Beyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Language

2025-10-27 · Mena Attia, Aashiq Muhamed, Mai Alkhamissi, Thamar Solorio 외 arxiv

We present a comprehensive evaluation of the ability of large language models (LLMs) to process culturally grounded language, specifically to understand and pragmatically use figurative expressions that encode local know…

English Proverbs

Catalog-Native LLM: Speaking Item-ID Dialect with Less Entanglement for Recommendation

2025-09-30 · Reza Shirkavand, Xiaokai Wei, Chen Wang, Zheng Hui 외 arxiv

While collaborative filtering delivers predictive accuracy and efficiency, and Large Language Models (LLMs) enable expressive and generalizable reasoning, modern recommendation systems must bring these strengths together…

Collaborative FilteringRecommendation Systems

A Parallel Cross-Lingual Benchmark for Multimodal Idiomaticity Understanding

2026-01-13 · Dilara Torunoğlu-Selamet, Dogukan Arslan, Rodrigo Wilkens, Wei He 외 arxiv

Potentially idiomatic expressions (PIEs) construe meanings inherently tied to the everyday experience of a given language community. As such, they constitute an interesting challenge for assessing the linguistic (and to …