paper-with-me

홈 › Papers

Following the Whispers of Values: Unraveling Neural Mechanisms Behind Value-Oriented Behaviors in LLMs

2025-04-07 · Ling Hu, Yuemei Xu, Xiaoyang Gu, Letao Han

Despite the impressive performance of large language models (LLMs), they can present unintended biases and harmful behaviors driven by encoded values, emphasizing the urgent need to understand the value mechanisms behind them. However, current research primarily evaluates these values through external responses with a focus on AI safety, lacking interpretability and failing to assess social values in real-world contexts. In this paper, we propose a novel framework called ValueExploration, which aims to explore the behavior-driven mechanisms of National Social Values within LLMs at the neuron level. As a case study, we focus on Chinese Social Values and first construct C-voice, a large-scale bilingual benchmark for identifying and evaluating Chinese Social Values in LLMs. By leveraging C-voice, we then identify and locate the neurons responsible for encoding these values according to activation difference. Finally, by deactivating these neurons, we analyze shifts in model behavior, uncovering the internal mechanism by which values influence LLM decision-making. Extensive experiments on four representative LLMs validate the efficacy of our framework. The benchmark and code will be available.

📄 PDF Abstract BibTeX arXiv:2504.04994

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Making

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Advancing NAM-to-Speech Conversion with Novel Methods and the MultiNAM Dataset

2024-12-25 · Neil Shah, Shirish Karande, Vineet Gandhi

Current Non-Audible Murmur (NAM)-to-speech techniques rely on voice cloning to simulate ground-truth speech from paired whispers. However, the simulated speech often lacks intelligibility and fails to generalize well acr…

text-to-speechText to SpeechVoice Cloning

Unraveling the Molecular Magic: AI Insights on the Formation of Extraordinarily Stretchable Hydrogels

2024-03-08 · Shahriar Hojjati Emmami, Ali Pilehvar Meibody, Lobat Tayebi, Mohammadamin Tavakoli 외

The deliberate manipulation of ammonium persulfate, methylenebisacrylamide, dimethyleacrylamide, and polyethylene oxide concentrations resulted in the development of a hydrogel with an exceptional stretchability, capable…

Chinese Whispers: A Multimodal Dataset for Embodied Language Grounding

2020-05-01 · LREC 2020 5 · Dimosthenis Kontogiorgos, Elena Sibirtseva, Joakim Gustafson

In this paper, we introduce a multimodal dataset in which subjects are instructing each other how to assemble IKEA furniture. Using the concept of {`}Chinese Whispers{'}, an old children{'}s game, we employ a novel metho…

Fourier Circuits in Neural Networks and Transformers: A Case Study of Modular Arithmetic with Multiple Inputs

2024-02-12 · Chenyang Li, YIngyu Liang, Zhenmei Shi, Zhao Song 외

In the evolving landscape of machine learning, a pivotal challenge lies in deciphering the internal representations harnessed by neural networks and Transformers. Building on recent progress toward comprehending how netw…

2kMathematical Reasoning

Heavy-Tailed Regularization of Weight Matrices in Deep Neural Networks

2023-04-06 · Xuanzhe Xiao, Zeng Li, Chuanlong Xie, Fengwei Zhou

Unraveling the reasons behind the remarkable success and exceptional generalization capabilities of deep neural networks presents a formidable challenge. Recent insights from random matrix theory, specifically those conc…