paper-with-me

홈 › Papers

Private Release of Text Embedding Vectors

2021-06-01 · NAACL (TrustNLP) 2021 6 · Oluwaseyi Feyisetan, Shiva Kasiviswanathan

Ensuring strong theoretical privacy guarantees on text data is a challenging problem which is usually attained at the expense of utility. However, to improve the practicality of privacy preserving text analyses, it is essential to design algorithms that better optimize this tradeoff. To address this challenge, we propose a release mechanism that takes any (text) embedding vector as input and releases a corresponding private vector. The mechanism satisfies an extension of differential privacy to metric spaces. Our idea based on first randomly projecting the vectors to a lower-dimensional space and then adding noise in this projected space generates private vectors that achieve strong theoretical guarantees on its utility. We support our theoretical proofs with empirical experiments on multiple word embedding models and NLP datasets, achieving in some cases more than 10% gains over the existing state-of-the-art privatization techniques.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Privacy Preserving

Similar Papers 제목 키워드 기반

Canonicalized Stable-List Replay for Private Federated Continual Learning over Language-Model Embeddings

2026-05-29 · Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Anuj Sharma arxiv

Federated continual learning (FCL) lets distributed clients adapt language-model heads to evolving NLP tasks without sharing raw text. Under user-level differential privacy (DP), replay-based continual learning faces a s…

Continual Learning

AdvSGM: Differentially Private Graph Learning via Adversarial Skip-gram Model

2025-03-27 · Sen Zhang, Qingqing Ye, Haibo Hu, Jianliang Xu

The skip-gram model (SGM), which employs a neural network to generate node vectors, serves as the basis for numerous popular graph embedding techniques. However, since the training datasets contain sensitive linkage info…

Graph EmbeddingGraph Learning

PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and Restoration

2024-06-03 · Ziqian Zeng, Jianwei Wang, Junyao Yang, Zhengdong Lu 외

The widespread usage of online Large Language Models (LLMs) inference services has raised significant privacy concerns about the potential exposure of private information in user inputs to malicious eavesdroppers. Existi…

Privacy Preserving

Word Embeddings for the Construction Domain

2016-10-28 · Antoine J. -P. Tixier, Michalis Vazirgiannis, Matthew R. Hallowell

We introduce word vectors for the construction domain. Our vectors were obtained by running word2vec on an 11M-word corpus that we created from scratch by leveraging freely-accessible online sources of construction-relat…

BenchmarkingGeneral ClassificationWord Embeddings

Privacy-Preserving Text Classification on BERT Embeddings with Homomorphic Encryption

2022-10-05 · NAACL 2022 7 · Garam Lee, Minsoo Kim, Jai Hyun Park, Seung-won Hwang 외

Embeddings, which compress information in raw text into semantics-preserving low-dimensional vectors, have been widely adopted for their efficacy. However, recent research has shown that embeddings can potentially leak p…

ClassificationGPUPrivacy Preservingtext-classification+1