paper-with-me

홈 › Papers

Differentially Private Steering for Large Language Model Alignment

2025-01-30 · Anmol Goel, Yaxi Hu, Iryna Gurevych, Amartya Sanyal

Aligning Large Language Models (LLMs) with human values and away from undesirable behaviors (such as hallucination) has become increasingly important. Recently, steering LLMs towards a desired behavior via activation editing has emerged as an effective method to mitigate harmful generations at inference-time. Activation editing modifies LLM representations by preserving information from positive demonstrations (e.g., truthful) and minimising information from negative demonstrations (e.g., hallucinations). When these demonstrations come from a private dataset, the aligned LLM may leak private information contained in those private samples. In this work, we present the first study of aligning LLM behavior with private datasets. Our work proposes the \textit{\underline{P}rivate \underline{S}teering for LLM \underline{A}lignment (PSA)} algorithm to edit LLM activations with differential privacy (DP) guarantees. We conduct extensive experiments on seven different benchmarks with open-source LLMs of different sizes (0.5B to 7B) and model families (LlaMa, Qwen, Mistral and Gemma). Our results show that PSA achieves DP guarantees for LLM alignment with minimal loss in performance, including alignment metrics, open-ended text generation quality, and general-purpose reasoning. We also develop the first Membership Inference Attack (MIA) for evaluating and auditing the empirical privacy for the problem of LLM steering via activation editing. Our attack is tailored for activation editing and relies solely on the generated texts without their associated probabilities. Our experiments support the theoretical guarantees by showing improved guarantees for our \textit{PSA} algorithm compared to several existing non-private techniques.

📄 PDF Abstract BibTeX arXiv:2501.18532

Code (1)

ukplab/iclr2025-psa 공식 구현 pytorch

Tasks

HallucinationInference AttackLanguage ModelingLanguage ModellingLarge Language ModelMembership Inference AttackmodelText Generation

Similar Papers 제목 키워드 기반

Differentially Private Preference Data Synthesis for Large Language Model Alignment

2026-05-29 · Fengyu Gao, Jing Yang arxiv

Preference alignment is a crucial post-training step for large language models (LLMs) to ensure their outputs align with human values. However, post-training on real human preference data raises privacy concerns, as thes…

EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors

2026-01-31 · Amin Banayeeanzade, Qingchuan Yang, Deqing Fu, Spencer Hong 외 arxiv

High-quality data is essential for modern machine learning, yet many valuable corpora are sensitive and cannot be freely shared. Synthetic data offers a practical substitute for downstream development, and large language…

Synthetic Data GenerationText Generation

Differentially Private Deep Learning with Direct Feedback Alignment

2020-10-08 · Jaewoo Lee, Daniel Kifer

Standard methods for differentially private training of deep neural networks replace back-propagated mini-batch gradients with biased and noisy approximations to the gradient. These modifications to training often result…

Deep LearningPrivacy Preserving

PROPS: Progressively Private Self-alignment of Large Language Models

2025-08-09 · Noel Teku, Fengwei Tian, Payel Bhattacharjee, Souradip Chakraborty 외 arxiv

Alignment is a key step in developing Large Language Models (LLMs) using human feedback to ensure adherence to human values and societal norms. Dependence on human feedback raises privacy concerns about how much a labele…

Privacy-Preserving Reinforcement Learning from Human Feedback via Decoupled Reward Modeling

2026-03-23 · Young Hyun Cho, Will Wei Sun arxiv

Preference-based fine-tuning has become an important component in training large language models, and the data used at this stage may contain sensitive user information. A central question is how to design a differential…

Reinforcement Learning