paper-with-me

홈 › Papers

Cultural Value Alignment Via Latent Activation Steering in Large Language Models

2026-05-25 · Trung Duc Anh Dang, Sarah Masud arxiv

Large Language Models (LLMs) often exhibit homogenized cultural perspectives. While the World Values Survey (WVS) provides a gold standard for mapping human values, traditional direct prompting of LLMs on WVS often fails to access the model's latent cultural depth, leading to safety-aligned refusals or neutral responses. Here, we propose a generalizable framework for cultural evaluation and intervention that transitions from abstract queries to scenario-based behavioral probing. By extracting implicit token probabilities across 300 situational dilemmas, we bypass surface-level alignment to map the latent coordinates of LLMs cultural value. We further introduce activation steering to shift these internal alignments during the forward pass without retraining. Across multiple LLMs, we find substantial variation in adaptability and uncover a consistent phenomenon of latent entanglement, where interventions along one cultural dimension induce shifts along another. These results suggest that cultural values are encoded as coupled structures, limiting precise alignment. This work establishes a computationally efficient framework for cultural steering, highlighting the structural complexities when navigating global value with LLMs.

📄 PDF Abstract BibTeX arXiv:2605.26365

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scenario-based Probing and Steering Cultural Values in Large Language Models--Extended Version

2026-06-09 · Trung Duc Anh Dang, Tung Kieu, Sarah Masud arxiv

Large Language Models (LLMs) are deployed across cultural contexts but often reflect homogenized values inherited from training data. Evaluations of cultural alignment typically rely on direct prompting with survey-style…

YaPO: Learnable Sparse Activation Steering Vectors for Domain Adaptation

2026-01-13 · Abdelaziz Bounhar, Rania Hossam Elmohamady Elbadry, Hadi Abdine, Preslav Nakov 외 arxiv

Steering Large Language Models (LLMs) through activation interventions has emerged as a lightweight alternative to fine-tuning for alignment and personalization. Recent work on Bi-directional Preference Optimization (BiP…

General KnowledgeDomain Adaptation

Steering Geometry: Validating Human Value Geometry in LLM Steering Space

2026-09-05 · Mohammad Mahdi Abootorabi, Armin Saghafian, Ali Bazshoushtari, Hamid Rezaei 외 hf

As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inference-time alternative to fine-tuning methods (e.g., RLHF, DPO) for behavio…

Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs

2026-05-07 · Andy Zeyi Liu, Michael Zhang, Ilana Greenberg, Adam Alnasser 외 arxiv

Steering large language models (LLMs) is usually done by either instruction prompting or activation steering. Prompting often gives strong control, but caches guidance tokens at every layer and can clutter long interacti…

DFKI-MLT at SemEval-2026 TASK 7: Steering Multilingual Models Towards Cultural Knowledge

2026-05-21 · Yusser Al Ghussin, Daniil Gurgurov, Yasser Hamidullah, Josef van Genabith 외 arxiv

Large language models (LLMs) are increasingly used across diverse linguistic and cultural contexts, yet their cultural knowledge remains uneven across regions and languages. We present the DFKI-MLT system for SemEval-202…