paper-with-me

홈 › Papers

Spectral Editing of Activations for Large Language Model Alignment

2024-05-15 · Yifu Qiu, Zheng Zhao, Yftah Ziser, Anna Korhonen, Edoardo M. Ponti, Shay B. Cohen

Large language models (LLMs) often exhibit undesirable behaviours, such as generating untruthful or biased content. Editing their internal representations has been shown to be effective in mitigating such behaviours on top of the existing alignment methods. We propose a novel inference-time editing method, namely spectral editing of activations (SEA), to project the input representations into directions with maximal covariance with the positive demonstrations (e.g., truthful) while minimising covariance with the negative demonstrations (e.g., hallucinated). We also extend our method to non-linear editing using feature functions. We run extensive experiments on benchmarks concerning truthfulness and bias with six open-source LLMs of different sizes and model families. The results demonstrate the superiority of SEA in effectiveness, generalisation to similar tasks, as well as computation and data efficiency. We also show that SEA editing only has a limited negative impact on other model capabilities.

📄 PDF Abstract BibTeX arXiv:2405.09719

Code (1)

yfqiu-nlp/sea-llm 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Differentially Private Steering for Large Language Model Alignment

2025-01-30 · Anmol Goel, Yaxi Hu, Iryna Gurevych, Amartya Sanyal

Aligning Large Language Models (LLMs) with human values and away from undesirable behaviors (such as hallucination) has become increasingly important. Recently, steering LLMs towards a desired behavior via activation edi…

HallucinationInference AttackLanguage ModelingLanguage Modelling+4

Why LoRA Fails to Forget: Regularized Low-Rank Adaptation Against Backdoors in Language Models

2026-01-09 · Hoang-Chau Luong, Lingwei Chen arxiv

Low-Rank Adaptation (LoRA) is widely used for parameter-efficient fine-tuning of large language models, but it is notably ineffective at removing backdoor behaviors from poisoned pretrained models when fine-tuning on cle…

parameter-efficient fine-tuning

Householder Pseudo-Rotation: A Novel Approach to Activation Editing in LLMs with Direction-Magnitude Perspective

2024-09-16 · Van-Cuong Pham, Thien Huu Nguyen

Activation Editing, which involves directly editting the internal representations of large language models (LLMs) to alter their behaviors and achieve desired properties, has emerged as a promising area of research. Exis…

AFTER: Mitigating the Object Hallucination of LVLM via Adaptive Factual-Guided Activation Editing

2026-01-05 · Tianbo Wang, Yuqing Ma, Kewei Liao, Zhange Zhang 외 arxiv

Large Vision-Language Models (LVLMs) have achieved substantial progress in cross-modal tasks. However, due to language bias, LVLMs are susceptible to object hallucination, which can be primarily divided into category, at…

WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language Models

2024-05-23 · Peng Wang, Zexi Li, Ningyu Zhang, Ziwen Xu 외

Large language models (LLMs) need knowledge updates to meet the ever-growing world facts and correct the hallucinated responses, facilitating the methods of lifelong model editing. Where the updated knowledge resides in …

HallucinationModel EditingQuestion AnsweringRetrieval