Interpretable Privacy Preservation of Text Representations Using Vector Steganography
Contextual word representations generated by language models (LMs) learn spurious associations present in the training corpora. Recent findings reveal that adversaries can exploit these associations to reverse-engineer the private attributes of entities mentioned within the corpora. These findings have led to efforts towards minimizing the privacy risks of language models. However, existing approaches lack interpretability, compromise on data utility and fail to provide privacy guarantees. Thus, the goal of my doctoral research is to develop interpretable approaches towards privacy preservation of text representations that retain data utility while guaranteeing privacy. To this end, I aim to study and develop methods to incorporate steganographic modifications within the vector geometry to obfuscate underlying spurious associations and preserve the distributional semantic properties learnt during training.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
On the Impact of Noise in Differentially Private Text Rewriting
The field of text privatization often leverages the notion of $\textit{Differential Privacy}$ (DP) to provide formal guarantees in the rewriting or obfuscation of sensitive textual data. A common and nearly ubiquitous fo…
SentenceUnveiling the Role of Message Passing in Dual-Privacy Preservation on GNNs
Graph Neural Networks (GNNs) are powerful tools for learning representations on graphs, such as social networks. However, their vulnerability to privacy inference attacks restricts their practicality, especially in high-…
Node ClassificationPrivacy PreservingNatural Language Understanding with Privacy-Preserving BERT
Privacy preservation remains a key challenge in data mining and Natural Language Understanding (NLU). Previous research shows that the input text or even text embeddings can leak private information. This concern motivat…
Language ModellingNatural Language UnderstandingPrivacy PreservingPrivAct: Internalizing Contextual Privacy Preservation via Multi-Agent Preference Training
Large language model (LLM) agents are increasingly deployed in personalized tasks involving sensitive, context-dependent information, where privacy violations may arise in agents' action due to the implicitness of contex…
Zero-shot GeneralizationProtecting Big Data Privacy Using Randomized Tensor Network Decomposition and Dispersed Tensor Computation
Data privacy is an important issue for organizations and enterprises to securely outsource data storage, sharing, and computation on clouds / fogs. However, data encryption is complicated in terms of the key management a…
Dimensionality ReductionManagementTensor Networks