paper-with-me

Papers

Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook

2026-03-16 · Jaehyeok Lee, Xiaoyuan Yi, Jing Yao, Hyunjin Hwang, Roy Ka-Wei Lee, Xing Xie, JinYeong Bak arxiv

As LLMs are globally deployed, aligning their cultural value orientations is critical for safety and user engagement. However, existing benchmarks face the Construct-Composition-Context ($C^3$) challenge: relying on discriminative, multiple-choice formats that probe value knowledge rather than true orientations, overlook subcultural heterogeneity, and mismatch with real-world open-ended generation. We introduce DOVE, a distributional evaluation framework that directly compares human-written text distributions with LLM-generated outputs. DOVE utilizes a rate-distortion variational optimization objective to construct a compact value codebook from 10K documents, mapping text into a structured value space to filter semantic noise. Alignment is measured using unbalanced optimal transport, capturing intra-cultural distributional structures and subgroup diversity. Experiments across 12 LLMs show that DOVE achieves superior predictive validity, attaining a 31.56% correlation with downstream tasks, while maintaining high reliability with as few as 500 samples per culture.

📄 PDF Abstract BibTeX arXiv:2604.06210

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output

2024-11-09 · Elise Karinshak, Amanda Hu, Kewen Kong, Vishwanatha Rao 외

Immense effort has been dedicated to minimizing the presence of harmful or biased generative content and better aligning AI output to human intention; however, research investigating the cultural values of LLMs is still …

Extrinsic Evaluation of Cultural Competence in Large Language Models

2024-06-17 · Shaily Bhatt, Fernando Diaz

Productive interactions between diverse users and language technologies require outputs from the latter to be culturally relevant and sensitive. Prior works have evaluated models' knowledge of cultural norms, values, and…

Open-Ended Question AnsweringQuestion AnsweringStory GenerationText Generation+1

CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis

2025-05-26 · Ruixiang Feng, Shen Gao, Xiuying Chen, Lisi Chen 외

Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, yet they often exhibit a specific cultural biases, neglecting the values and linguistic diversity of low-resource regions. This…

DiversityOpen-Ended Question AnsweringQuestion Answering

Scenario-based Probing and Steering Cultural Values in Large Language Models--Extended Version

2026-06-09 · Trung Duc Anh Dang, Tung Kieu, Sarah Masud arxiv

Large Language Models (LLMs) are deployed across cultural contexts but often reflect homogenized values inherited from training data. Evaluations of cultural alignment typically rely on direct prompting with survey-style…

From National Curricula to Cultural Awareness: Constructing Open-Ended Culture-Specific Question Answering Dataset

2026-01-08 · Haneul Yoo, Won Ik Cho, Geunhye Kim, Jiyoon Han arxiv

Large language models (LLMs) achieve strong performance on many tasks, but their progress remains uneven across languages and cultures, often reflecting values latent in English-centric training data. To enable practical…

Question Answering