paper-with-me

Papers

Do Large Language Models Truly Understand Cross-cultural Differences?

2025-12-08 · Shiwei Guo, Sihang Jiang, Qianxi He, Yanghua Xiao, Jiaqing Liang, Bi Yude, Minggui He, Shimin Tao, Li Zhang arxiv

In recent years, large language models (LLMs) have demonstrated strong performance on multilingual tasks. Given its wide range of applications, cross-cultural understanding capability is a crucial competency. However, existing benchmarks for evaluating whether LLMs genuinely possess this capability suffer from three key limitations: a lack of contextual scenarios, insufficient cross-cultural concept mapping, and limited deep cultural reasoning capabilities. To address these gaps, we propose SAGE, a scenario-based benchmark built via cross-cultural core concept alignment and generative task design, to evaluate LLMs' cross-cultural understanding and reasoning. Grounded in cultural theory, we categorize cross-cultural capabilities into nine dimensions. Using this framework, we curated 210 core concepts and constructed 4530 test items across 15 specific real-world scenarios, organized under four broader categories of cross-cultural situations, following established item design principles. The SAGE dataset supports continuous expansion, and experiments confirm its transferability to other languages. It reveals model weaknesses across both dimensions and scenarios, exposing systematic limitations in cross-cultural reasoning. While progress has been made, LLMs are still some distance away from reaching a truly nuanced cross-cultural understanding. In compliance with the anonymity policy, we include data and code in the supplement materials. In future versions, we will make them publicly available online.

📄 PDF Abstract BibTeX arXiv:2512.07075

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Do LLMs Capture Embodied Cognition and Cultural Variation? Cross-Linguistic Evidence from Demonstratives

2026-04-28 · Yu Wang, Emmanuele Chersoni, Chu-Ren Huang arxiv

Do large language models (LLMs) truly acquire embodied cognition and cultural conventions from text? We introduce demonstratives, fundamental spatial expressions like "this/that" in English and "zhè/nà" in Chinese, as a …

Metaphors We Compute By: A Computational Audit of Cultural Translation vs. Thinking in LLMs

2026-04-06 · Yuan Chang, Jiaming Qu, Zhu Li arxiv

Large language models (LLMs) are often described as multilingual because they can understand and respond in many languages. However, speaking a language is not the same as reasoning within a culture. This distinction mot…

Fluent but Culturally Distant: Can Regional Training Teach Cultural Understanding?

2025-05-25 · Dhruv Agarwal, Anya Shukla, Sunayana Sitaram, Aditya Vashistha

Large language models (LLMs) are used around the world but exhibit Western cultural tendencies. To address this cultural misalignment, many countries have begun developing "regional" LLMs tailored to local communities. Y…

Everyday Physics in Korean Contexts: A Culturally Grounded Physical Reasoning Benchmark

2025-09-22 · Jihae Jeong, DaeYeop Lee, DongGeon Lee, Hwanjo Yu arxiv

Existing physical commonsense reasoning benchmarks predominantly focus on Western contexts, overlooking cultural variations in physical problem-solving. To address this gap, we introduce EPiK (Everyday Physics in Korean …

Physical Commonsense Reasoning

Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs

2025-05-21 · Hao Wang, Pinzhi Huang, Jihan Yang, Saining Xie 외

The rapid evolution of multimodal large language models (MLLMs) has significantly enhanced their real-world applications. However, achieving consistent performance across languages, especially when integrating cultural k…

BenchmarkingQuestion AnsweringVisual Question Answering