paper-with-me

홈 › Papers

Vision Language Model Helps Private Information De-Identification in Vision Data

2026-06-08 · Tiejin Chen, Pingzhi Li, Kaixiong Zhou, Tianlong Chen, Hua Wei arxiv

Visual Language Models (VLMs) have gained significant popularity due to their remarkable ability. While various methods exist to enhance privacy in text-based applications, privacy risks associated with visual inputs remain largely overlooked such as Protected Health Information (PHI) in medical images. To tackle this problem, two key tasks: accurately localizing sensitive text and processing it to ensure privacy protection should be performed. To address this issue, we introduce VisShield (Vision Privacy Shield), an end-to-end framework designed to enhance the privacy awareness of VLMs. Our framework consists of two key components: a specialized instruction-tuning dataset OPTIC (Optical Privacy Text Instruction Collection) and a tailored training methodology. The dataset provides diverse privacy-oriented prompts that guide VLMs to perform targeted Optical Character Recognition (OCR) for precise localization of sensitive text, while the training strategy ensures effective adaptation of VLMs to privacy-preserving tasks. Specifically, our approach ensures that VLMs recognize privacy-sensitive text and output precise bounding boxes for detected entities, allowing for effective masking of sensitive information. Extensive experiments demonstrate that our framework significantly outperforms existing approaches in handling private information, paving the way for privacy-preserving applications in vision-language models. Our dataset and code can be found here.

📄 PDF Abstract BibTeX arXiv:2606.09132

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Coupling public and private gradient provably helps optimization

2023-10-02 · Ruixuan Liu, Zhiqi Bu, Yu-Xiang Wang, Sheng Zha 외

The success of large neural networks is crucially determined by the availability of data. It has been observed that training only on a small amount of public data, or privately on the abundant private data can lead to un…

Differentially Private Language Generation and Identification in the Limit

2026-04-09 · Anay Mehrotra, Grigoris Velegkas, Xifan Yu, Felix Zhou arxiv

We initiate the study of language generation in the limit, a model recently introduced by Kleinberg and Mullainathan [KM24], under the constraint of differential privacy. We consider the continual release model, where a …

Language Identification

An Easy-to-use and Robust Approach for the Differentially Private De-Identification of Clinical Textual Documents

2022-11-02 · Yakini Tchouka, Jean-François Couchot, David Laiymani

Unstructured textual data is at the heart of healthcare systems. For obvious privacy reasons, these documents are not accessible to researchers as long as they contain personally identifiable information. One way to shar…

De-identificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

Mobile Sensor Data Anonymization

2018-10-26 · Mohammad Malekzadeh, Richard G. Clegg, Andrea Cavallaro, Hamed Haddadi

Motion sensors such as accelerometers and gyroscopes measure the instant acceleration and rotation of a device, in three dimensions. Raw data streams from motion sensors embedded in portable and wearable devices may reve…

Activity RecognitionDecoderUser Identification

A survey on phrase structure learning methods for text classification

2014-06-21 · Reshma Prasad, Mary Priya Sebastian

Text classification is a task of automatic classification of text into one of the predefined categories. The problem of text classification has been widely studied in different communities like natural language processin…

ClassificationGeneral ClassificationGenre classificationInformation Retrieval+8