paper-with-me

홈 › Papers

CARTE: A Benchmark for Mapping Language Model Knowledge Across France

2026-06-01 · Sarah Almeida Carneiro, Christos Xypolopoulos, Xiao Fei, Yang Zhang, Michalis Vazirgiannis arxiv

We introduce CARTE 1 (Culturally Anchored Regional-Territorial Evaluation), a multiplechoice benchmark for evaluating the ability of large language models (LLMs) to perform fine-grained reasoning over geographically grounded and regionally differentiated knowledge within France. While prior benchmarks focus on national-level cultural understanding, they largely overlook intra-country variation and the need to distinguish between closely related regional contexts. CARTE addresses this gap by introducing 2,431 questions spanning the 13 metropolitan regions of France and covering 14 thematic domains, including culture, language, demographics, economy, environment, and mobility. We further introduce CARTE-LV, a subset targeting Linguistic Variation across French regions, enabling focused evaluation of language-related differences. We evaluate 27 LLMs ranging from 1B to 12B parameters under few-shot settings. Our experiments reveal performance disparities across regions and model scales, suggesting systematic gaps in pretraining coverage and limited robustness to intra-national variation.

📄 PDF Abstract BibTeX arXiv:2606.01995

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Descartes: Generating Short Descriptions of Wikipedia Articles

2022-05-20 · Marija Sakota, Maxime Peyrard, Robert West

Wikipedia is one of the richest knowledge sources on the Web today. In order to facilitate navigating, searching, and maintaining its content, Wikipedia's guidelines state that all articles should be annotated with a so-…

Articles

LAG: Logic-Augmented Generation from a Cartesian Perspective

2025-08-07 · Yilin Xiao, Chuang Zhou, Yujing Zhang, Qinggang Zhang 외 arxiv

Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks, yet exhibit critical limitations in knowledge-intensive tasks, often generating hallucinations when faced with question…

Semantic Retrieval

CartesianMoE: Boosting Knowledge Sharing among Experts via Cartesian Product Routing in Mixture-of-Experts

2024-10-21 · Zhenpeng Su, Xing Wu, Zijia Lin, Yizhe Xiong 외

Large language models (LLM) have been attracting much attention from the community recently, due to their remarkable performance in all kinds of downstream tasks. According to the well-known scaling law, scaling up a den…

Mixture-of-Experts

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space

2026-05-11 · Xia Hu, Zhenrui Yue, Brian Potetz, Howard Zhou 외 arxiv

As current Multimodal Large Language Models rapidly saturate canonical visual reasoning benchmarks, a key question emerges: do these strong scores genuinely reflect robust visual understanding? We identify a pervasive vu…

Visual Reasoning

Integrating Machine-Generated Short Descriptions into the Wikipedia Android App: A Pilot Deployment of Descartes

2026-01-12 · Marija Šakota, Dmitry Brant, Cooltey Feng, Shay Nowick 외 arxiv

Short descriptions are a key part of the Wikipedia user experience, but their coverage remains uneven across languages and topics. In previous work, we introduced Descartes, a multilingual model for generating short desc…