paper-with-me

홈 › Papers

Scenario-based Probing and Steering Cultural Values in Large Language Models--Extended Version

2026-06-09 · Trung Duc Anh Dang, Tung Kieu, Sarah Masud arxiv

Large Language Models (LLMs) are deployed across cultural contexts but often reflect homogenized values inherited from training data. Evaluations of cultural alignment typically rely on direct prompting with survey-style questions, which frequently elicit neutral or safety-aligned responses and fail to capture underlying model preferences. We propose a framework for probing and steering latent cultural representations in LLMs along the two Inglehart--Welzel axes of the World Values Survey (WVS). By translating social value questions into scenario-based behavioral dilemmas, we extract token-level probabilities to measure implicit values and apply activation steering, optionally combined with country-conditioned prompting, to shift model behavior without retraining. Across three open-source LLMs and four target cultures, we find substantial variation in steerability and identify latent entanglement, where interventions along one cultural dimension induce shifts along another. This coupling mirrors correlations in human WVS data and persists across activation, prompt, and hybrid steering. It constrains axis-independent alignment, though general task performance is largely preserved.

📄 PDF Abstract BibTeX arXiv:2606.11399

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Cultural Value Alignment Via Latent Activation Steering in Large Language Models

2026-05-25 · Trung Duc Anh Dang, Sarah Masud arxiv

Large Language Models (LLMs) often exhibit homogenized cultural perspectives. While the World Values Survey (WVS) provides a gold standard for mapping human values, traditional direct prompting of LLMs on WVS often fails…

Training-Free Cultural Alignment of Large Language Models via Persona Disagreement

2026-05-11 · Huynh Trung Kiet, Dao Sy Duy Minh, Tuan Nguyen, Chi-Nguyen Tran 외 arxiv

Large language models increasingly mediate decisions that turn on moral judgement, yet a growing body of evidence shows that their implicit preferences are not culturally neutral. Existing cultural alignment methods eith…

Probing Pre-Trained Language Models for Cross-Cultural Differences in Values

2022-03-25 · Arnav Arora, Lucie-Aimée Kaffee, Isabelle Augenstein

Language embeds information about social, cultural, and political values people hold. Prior work has explored social and potentially harmful biases encoded in Pre-Trained Language models (PTLMs). However, there has been …

Cultural Binding Heads in Language Models

2026-05-27 · Avrile Floro, Luca Benedetto arxiv

LLMs often default to equal treatment across cultural groups, even though context warrants differentiation: this is a lack of difference awareness. Using mechanistic interpretability and a factorial design on the N4 cult…

BACH-V: Bridging Abstract and Concrete Human-Values in Large Language Models

2026-01-20 · Junyu Zhang, Yipeng Kang, Jiong Guo, Jiayu Zhan 외 arxiv

Do large language models (LLMs) genuinely understand abstract concepts, or merely manipulate them as statistical patterns? We introduce an abstraction-grounding framework that decomposes conceptual understanding into thr…