paper-with-me

홈 › Papers

CARE: Aligning Language Models for Regional Cultural Awareness

2025-04-07 · Geyang Guo, Tarek Naous, Hiromi Wakaki, Yukiko Nishimura, Yuki Mitsufuji, Alan Ritter, Wei Xu

Existing language models (LMs) often exhibit a Western-centric bias and struggle to represent diverse cultural knowledge. Previous attempts to address this rely on synthetic data and express cultural knowledge only in English. In this work, we study whether a small amount of human-written, multilingual cultural preference data can improve LMs across various model families and sizes. We first introduce CARE, a multilingual resource of 24.1k responses with human preferences on 2,580 questions about Chinese and Arab cultures, all carefully annotated by native speakers and offering more balanced coverage. Using CARE, we demonstrate that cultural alignment improves existing LMs beyond generic resources without compromising general capabilities. Moreover, we evaluate the cultural awareness of LMs, native speakers, and retrieved web content when queried in different languages. Our experiment reveals regional disparities among LMs, which may also be reflected in the documentation gap: native speakers often take everyday cultural commonsense and social norms for granted, while non-natives are more likely to actively seek out and document them. CARE is publicly available at https://github.com/Guochry/CARE (we plan to add Japanese data in the near future).

📄 PDF Abstract BibTeX arXiv:2504.05154

Code (1)

guochry/care 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Evaluating Cultural Awareness of LLMs for Yoruba, Malayalam, and English

2024-09-14 · Fiifi Dawson, Zainab Mosunmola, Sahil Pocker, Raj Abhijit Dandekar 외

Although LLMs have been extremely effective in a large number of complex tasks, their understanding and functionality for regional languages and cultures are not well studied. In this paper, we explore the ability of var…

Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment

2025-09-26 · Hongbin Zhang, Kehai Chen, Xuefeng Bai, Yang Xiang 외 arxiv

Reward models (RMs) are crucial for aligning large language models (LLMs) with diverse cultures. Consequently, evaluating their cultural awareness is essential for further advancing global alignment of LLMs. However, exi…

Reinforcement Learning

DIWALI: Diversity and Inclusivity aWare cuLture specific Items for India: Dataset and Assessment of LLMs for Cultural Text Adaptation in Indian Context

2025-09-22 · Pramit Sahoo, Maharaj Brahma, Maunendra Sankar Desarkar arxiv

Large language models (LLMs) are widely used in various tasks and applications. However, despite their wide capabilities, they are shown to lack cultural alignment \citep{ryan-etal-2024-unintended, alkhamissi-etal-2024-i…

XLQA: A Benchmark for Locale-Aware Multilingual Open-Domain Question Answering

2025-08-22 · Keon-Woo Roh, Yeong-Joon Ju, Seong-Whan Lee arxiv

Large Language Models (LLMs) have shown significant progress in Open-domain question answering (ODQA), yet most evaluations focus on English and assume locale-invariant answers across languages. This assumption neglects …

Open-Domain Question Answering

Preserving Cultural Identity with Context-Aware Translation Through Multi-Agent AI Systems

2025-03-05 · Mahfuz Ahmed Anik, Abdur Rahman, Azmine Toushik Wasi, Md Manjurul Ahsan

Language is a cornerstone of cultural identity, yet globalization and the dominance of major languages have placed nearly 3,000 languages at risk of extinction. Existing AI-driven translation models prioritize efficiency…

DiversityTranslation