paper-with-me

홈 › Papers

The Cultural Gene of Large Language Models: A Study on the Impact of Cross-Corpus Training on Model Values and Biases

2025-08-17 · Emanuel Z. Fenech-Borg, Tilen P. Meznaric-Kos, Milica D. Lekovic-Bojovic, Arni J. Hentze-Djurhuus arxiv

Large language models (LLMs) are deployed globally, yet their underlying cultural and ethical assumptions remain underexplored. We propose the notion of a "cultural gene" -- a systematic value orientation that LLMs inherit from their training corpora -- and introduce a Cultural Probe Dataset (CPD) of 200 prompts targeting two classic cross-cultural dimensions: Individualism-Collectivism (IDV) and Power Distance (PDI). Using standardized zero-shot prompts, we compare a Western-centric model (GPT-4) and an Eastern-centric model (ERNIE Bot). Human annotation shows significant and consistent divergence across both dimensions. GPT-4 exhibits individualistic and low-power-distance tendencies (IDV score approx 1.21; PDI score approx -1.05), while ERNIE Bot shows collectivistic and higher-power-distance tendencies (IDV approx -0.89; PDI approx 0.76); differences are statistically significant (p < 0.001). We further compute a Cultural Alignment Index (CAI) against Hofstede's national scores and find GPT-4 aligns more closely with the USA (e.g., IDV CAI approx 0.91; PDI CAI approx 0.88) whereas ERNIE Bot aligns more closely with China (IDV CAI approx 0.85; PDI CAI approx 0.81). Qualitative analyses of dilemma resolution and authority-related judgments illustrate how these orientations surface in reasoning. Our results support the view that LLMs function as statistical mirrors of their cultural corpora and motivate culturally aware evaluation and deployment to avoid algorithmic cultural hegemony.

📄 PDF Abstract BibTeX arXiv:2508.12411

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense

2024-05-07 · Siqi Shen, Lajanugen Logeswaran, Moontae Lee, Honglak Lee 외

Large language models (LLMs) have demonstrated substantial commonsense understanding through numerous benchmark evaluations. However, their understanding of cultural commonsense remains largely unexamined. In this paper,…

Cultural Value Differences of LLMs: Prompt, Language, and Model Size

2024-06-17 · Qishuai Zhong, Yike Yun, Aixin Sun

Our study aims to identify behavior patterns in cultural values exhibited by large language models (LLMs). The studied variants include question ordering, prompting language, and model size. Our experiments reveal that e…

Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance

2024-02-22 · Ziqi Yin, Hao Wang, Kaito Horio, Daisuke Kawahara 외

We investigate the impact of politeness levels in prompts on the performance of large language models (LLMs). Polite language in human communications often garners more compliance and effectiveness, while rudeness can ca…

Risks of Cultural Erasure in Large Language Models

2025-01-02 · Rida Qadri, Aida M. Davani, Kevin Robinson, Vinodkumar Prabhakaran

Large language models are increasingly being integrated into applications that shape the production and discovery of societal knowledge such as search, online education, and travel planning. As a result, language models …

Language ModelingLanguage Modelling

Cross-Cultural Transfer Learning for Chinese Offensive Language Detection

2023-03-31 · Li Zhou, Laura Cabello, Yong Cao, Daniel Hershcovich

Detecting offensive language is a challenging task. Generalizing across different cultures and languages becomes even more challenging: besides lexical, syntactic and semantic differences, pragmatic aspects such as cultu…

Cultural Vocal Bursts Intensity PredictionFew-Shot LearningTransfer Learning