paper-with-me

홈 › Papers

Discovering Universal Activation Directions for PII Leakage in Language Models

2026-02-19 · Leo Marchyok, Zachary Coalson, Sungho Keum, Sooel Son, Sanghyun Hong arxiv

Modern language models exhibit rich internal structure, yet little is known about how privacy-sensitive behaviors, such as personally identifiable information (PII) leakage, are represented and modulated within their hidden states. We present UniLeak, a mechanistic-interpretability framework that identifies universal activation directions: latent directions in a model's residual stream whose linear addition at inference time consistently increases the likelihood of generating PII across prompts. These model-specific directions generalize across contexts and amplify PII generation probability, with minimal impact on generation quality. UniLeak recovers such directions without access to training data or groundtruth PII, relying only on self-generated text. Across multiple models and datasets, steering along these universal directions substantially increases PII leakage compared to existing prompt-based extraction methods. Our results offer a new perspective on PII leakage: the superposition of a latent signal in the model's representations, enabling both risk amplification and mitigation.

📄 PDF Abstract BibTeX arXiv:2602.16980

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Testing the Limits of Truth Directions in LLMs

2026-04-04 · Angelos Poulis, Mark Crovella, Evimaria Terzi arxiv

Large language models (LLMs) have been shown to encode truth of statements in their activation space along a linear truth direction. Previous studies have argued that these directions are universal in certain aspects, wh…

Zero-Direction Probing: A Linear-Algebraic Framework for Deep Analysis of Large-Language-Model Drift

2025-08-09 · Amit Pandey arxiv

We present Zero-Direction Probing (ZDP), a theory-only framework for detecting model drift from null directions of transformer activations without task labels or output evaluations. Under assumptions A1--A6, we prove: (i…

What Does the Server See? Understanding Privacy Leakage from Large Language Models in Split Inference

2026-05-22 · Mingyuan Fan, Yu Liu, Fuyi Wang, Cen Chen arxiv

The deployment of large language models (LLMs) on resource-constrained devices remains challenging, spurring interest in split inference, where models are partitioned between client and server to reduce computational bur…

Language Steering for Multilingual In-Context Learning

2026-02-02 · Neeraja Kirtane, Kuan-Hao Huang arxiv

If large language models operate in a universal semantic space, then switching between languages should require only a simple activation offset. To test this, we take multilingual in-context learning as a case study, whe…

Discovering Interpretable Directions in the Semantic Latent Space of Diffusion Models

2023-03-20 · René Haas, Inbar Huberman-Spiegelglas, Rotem Mulayoff, Stella Graßhof 외

Denoising Diffusion Models (DDMs) have emerged as a strong competitor to Generative Adversarial Networks (GANs). However, despite their widespread use in image synthesis and editing applications, their latent space is st…

AttributeDenoisingImage Generation