paper-with-me

Papers

Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations

2026-03-31 · Haoran Wang, Li Xiong, Kai Shu arxiv

Large language models (LLMs) are increasingly deployed in high-stakes settings, yet they frequently violate contextual privacy by disclosing private information in situations where humans would exercise discretion. This raises a fundamental question: do LLMs internally encode contextual privacy norms, and if so, why do violations persist? We present the first systematic study of contextual privacy as a structured latent representation in LLMs, grounded in contextual integrity (CI) theory. Probing multiple models, we find that the three norm-determining CI parameters (information type, recipient, and transmission principle) are encoded as linearly separable and functionally independent directions in activation space. Despite this internal structure, models still leak private information in practice, revealing a clear gap between concept representation and model behavior. To bridge this gap, we introduce CI-parametric steering, which independently intervenes along each CI dimension. This structured control reduces privacy violations more effectively and predictably than monolithic steering. Our results demonstrate that contextual privacy failures arise from misalignment between representation and behavior rather than missing awareness, and that leveraging the compositional structure of CI enables more reliable contextual privacy control, shedding light on potential improvement of contextual privacy understanding in LLMs.

📄 PDF Abstract BibTeX arXiv:2604.00209

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Analysing the Residual Stream of Language Models Under Knowledge Conflicts

2024-10-21 · Yu Zhao, Xiaotang Du, Giwon Hong, Aryo Pradipta Gema 외

Large language models (LLMs) can store a significant amount of factual knowledge in their parameters. However, their parametric knowledge may conflict with the information provided in the context. Such conflicts can lead…

Behavior and Representation in Large Language Models for Combinatorial Optimization: From Feature Extraction to Algorithm Selection

2025-12-15 · Francesca Da Ros, Luca Di Gaspero, Kevin Roitero arxiv

Recent advances in Large Language Models (LLMs) have opened new perspectives for automation in optimization. While several studies have explored how LLMs can generate or solve optimization models, far less is understood …

Detecting Is Not Resolving: The Monitoring Control Gap in Retrieval Augmented LLMs

2026-05-26 · Zhe Yu, Wenpeng Xing, Chen Ye, Xuyang Teng 외 arxiv

Retrieval-augmented LLMs are deployed for tasks where evidence quality determines action safety, yet evaluation protocols assume that single-turn robustness predicts robustness when evidence accumulates across turns. We …

Mapping Clinical Doubt: Locating Linguistic Uncertainty in LLMs

2025-11-27 · Srivarshinee Sridhar, Raghav Kaushik Ravi, Kripabandhu Ghosh arxiv

Large Language Models (LLMs) are increasingly used in clinical settings, where sensitivity to linguistic uncertainty can influence diagnostic interpretation and decision-making. Yet little is known about where such epist…

Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution

2024-09-30 · Haiyan Zhao, Heng Zhao, Bo Shen, Ali Payani 외

Probing learned concepts in large language models (LLMs) is crucial for understanding how semantic knowledge is encoded internally. Training linear classifiers on probing tasks is a principle approach to denote the vecto…

Text Generation