paper-with-me

홈 › Papers

DPN-LE: Dual Personality Neuron Localization and Editing for Large Language Models

2026-04-30 · Lifan Zheng, Xue Yang, Jiawei Chen, Chenyan Wu, Jingyuan Zhang, Fanheng Kong, Xinyi Zeng, Xiang Chen, Yu Tian arxiv

With the widespread adoption of large language models (LLMs), understanding their personality representation mechanisms has become critical. As a novel paradigm in Personality Editing, most existing methods employ neuron-editing to locate and modify LLM neurons, requiring changes to numerous neurons and leading to significant performance degradation. This raises a fundamental question: Are all modified neurons directly related to personality representation? In this work, we investigate and quantify this specificity through assessments of general capability impact and representation-level patterns. We find that: 1) Current methods can change personalities but reduce overall performance. 2) Neurons are multifunctional, connecting personality traits and general knowledge. 3) Opposing personality traits demonstrate distinctly mutually exclusive representation patterns. Motivated by these findings, we propose DPN-LE (Dual Personality Neuron Localization and Editing), which identifies personality-specific neurons by contrasting MLP activations between high-trait and low-trait samples. DPN-LE constructs layer-wise steering vectors and applies dual-criterion filtering based on Cohen's $d$ effect size and activation magnitude to isolate mutually exclusive neuron subsets. Sparse linear intervention on these neurons enables precise personality control at inference time. Using only 1,000 contrastive sample pairs per trait, DPN-LE intervenes on $\sim$0.5\% of neurons while achieving competitive personality control and substantially better capability preservation across reasoning tasks. Experiments on LLaMA-3-8B-Instruct and Qwen2.5-7B-Instruct demonstrate the effectiveness and generalizability of our approach.

📄 PDF Abstract BibTeX arXiv:2604.27929

Code (0)

등록된 구현이 없습니다.

Tasks

General Knowledge

Similar Papers 제목 키워드 기반

Knowledge Editing for Large Language Model with Knowledge Neuronal Ensemble

2024-12-30 · Yongchang Li, Yujin Zhu, Tao Yan, Shijian Fan 외

As real-world knowledge is constantly evolving, ensuring the timeliness and accuracy of a model's knowledge is crucial. This has made knowledge editing in large language models increasingly important. However, existing k…

knowledge editingLanguage ModelingLanguage ModellingLarge Language Model+1

Editing Personality for Large Language Models

2023-10-03 · Shengyu Mao, Xiaohan Wang, Mengru Wang, Yong Jiang 외

This paper introduces an innovative task focused on editing the personality traits of Large Language Models (LLMs). This task seeks to adjust the models' responses to opinion-related questions on specified topics since a…

Model Editing

Precise Localization of Memories: A Fine-grained Neuron-level Knowledge Editing Technique for LLMs

2025-03-03 · Haowen Pan, Xiaozhi Wang, Yixin Cao, Zenglin Shi 외

Knowledge editing aims to update outdated information in Large Language Models (LLMs). A representative line of study is locate-then-edit methods, which typically employ causal tracing to identify the modules responsible…

knowledge editing

Neuron-based Personality Trait Induction in Large Language Models

2024-10-16 · Jia Deng, Tianyi Tang, Yanbin Yin, Wenhao Yang 외

Large language models (LLMs) have become increasingly proficient at simulating various personality traits, an important capability for supporting related applications (e.g., role-playing). To further improve this capacit…

Output Vector Editing for Memorization Mitigation in Large Language Models

2026-06-17 · Ahmad Dawar Hakimi, Kaiwei Lei, Isabelle Augenstein, Hinrich Schütze arxiv

Large language models memorize and reproduce sequences from their training data, creating privacy, copyright, and security risks. Existing neuron-level mitigation methods equate editing with zeroing out neuron activation…