paper-with-me

홈 › Papers

Vectors Are Not Neutral: Sensitive-Information Inference from Exported LLM Representations in Summarization

2026-05-26 · Weixin Liu, Bowen Qu, Juming Xiong, Congning Ni, Bradley A. Malin, Zhijun Yin arxiv

Large language model (LLM) summarization systems may pass compact vector representations of private inputs to downstream retrieval, monitoring, audit, or analytic workflows. Even when source documents remain access-restricted, derived vectors may be handled under different access controls and still support sensitive-information inference, creating a residual information-disclosure risk. We study this issue in clinical discharge-summary generation as a high-stakes case study, using electronic health record (EHR)-recorded race as a controlled sensitive-label audit. We audit two artifacts that a system might retain or expose to downstream components: the final prompt-token hidden state and the mean-pooled prompt representation. Our results show that reducing recoverability of the case-study sensitive label from one exported artifact does not necessarily reduce recoverability from another. As a mitigation case study, we introduce SurfaceLoRA, an exported-vector-targeted parameter-efficient fine-tuning method that uses a gradient-reversal discriminator attached to a designated exported vector. Under a balanced five-way probing protocol, SurfaceLoRA reduces EHR-recorded race recoverability from the targeted final-token artifact toward chance while preserving summarization utility, yet recoverability remains substantially higher from untargeted pooled artifacts. These findings show that privacy auditing and mitigation should be performed on the exact vector artifact retained or exposed to downstream components.

📄 PDF Abstract BibTeX arXiv:2605.26433

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuning

Similar Papers 제목 키워드 기반

Informationally Compressive Anonymization: Non-Degrading Sensitive Input Protection for Privacy-Preserving Supervised Machine Learning

2026-03-16 · Jeremy J Samuelson arxiv

Modern machine learning systems increasingly rely on sensitive data, creating significant privacy, security, and regulatory risks that existing privacy-preserving machine learning (ppML) techniques, such as Differential …

Representation Learning

FairSIN: Achieving Fairness in Graph Neural Networks through Sensitive Information Neutralization

2024-03-19 · Cheng Yang, Jixi Liu, Yunhe Yan, Chuan Shi

Despite the remarkable success of graph neural networks (GNNs) in modeling graph-structured data, like other machine learning models, GNNs are also susceptible to making biased predictions based on sensitive attributes, …

Fairness

The Geometry of Multilingual Language Model Representations

2022-05-22 · Tyler A. Chang, Zhuowen Tu, Benjamin K. Bergen

We assess how multilingual language models maintain a shared multilingual representation space while still encoding language-sensitive information in each language. Using XLM-R as a case study, we show that languages occ…

Cross-Lingual TransferLanguage ModelingLanguage Modellingmodel+2

Controllable Affective Generation via Latent Vector Steering

2026-08-26 · Xixian Yong, Siyuan Chang, Yingying Zhang, Xian Wu 외 arxiv

Large Language Models (LLMs) often produce emotionally flattened responses after alignment, limiting their effectiveness in affect-sensitive applications. In this paper, we propose EmoVec, a lightweight framework for con…

Continuous Control

Fairness via Representation Neutralization

2021-06-23 · NeurIPS 2021 12 · Mengnan Du, Subhabrata Mukherjee, Guanchu Wang, Ruixiang Tang 외

Existing bias mitigation methods for DNN models primarily work on learning debiased encoders. This process not only requires a lot of instance-level annotations for sensitive attributes, it also does not guarantee that a…

AttributeClassificationFairness