paper-with-me

홈 › Papers

Unveiling the Unseen: Exploring Whitebox Membership Inference through the Lens of Explainability

2024-07-01 · Chenxi Li, Abhinav Kumar, Zhen Guo, Jie Hou, Reza Tourani

The increasing prominence of deep learning applications and reliance on personalized data underscore the urgent need to address privacy vulnerabilities, particularly Membership Inference Attacks (MIAs). Despite numerous MIA studies, significant knowledge gaps persist, particularly regarding the impact of hidden features (in isolation) on attack efficacy and insufficient justification for the root causes of attacks based on raw data features. In this paper, we aim to address these knowledge gaps by first exploring statistical approaches to identify the most informative neurons and quantifying the significance of the hidden activations from the selected neurons on attack accuracy, in isolation and combination. Additionally, we propose an attack-driven explainable framework by integrating the target and attack models to identify the most influential features of raw data that lead to successful membership inference attacks. Our proposed MIA shows an improvement of up to 26% on state-of-the-art MIA.

📄 PDF Abstract BibTeX arXiv:2407.01306

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unveiling Impact of Frequency Components on Membership Inference Attacks for Diffusion Models

2025-05-27 · Puwei Lian, Yujun Cai, Songze Li

Diffusion models have achieved tremendous success in image generation, but they also raise significant concerns regarding privacy and copyright issues. Membership Inference Attacks (MIAs) are designed to ascertain whethe…

Image Generation

May I have your Attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks

2025-07-10 · Nishit V. Pandya, Andrey Labunets, Sicun Gao, Earlence Fernandes arxiv

A popular class of defenses against prompt injection attacks on large language models (LLMs) relies on fine-tuning to separate instructions and data, so that the LLM does not follow instructions that might be present wit…

Unveiling Client Privacy Leakage from Public Dataset Usage in Federated Distillation

2025-02-11 · Haonan Shi, Tu Ouyang, An Wang

Federated Distillation (FD) has emerged as a popular federated training framework, enabling clients to collaboratively train models without sharing private data. Public Dataset-Assisted Federated Distillation (PDA-FD), w…

Federated LearningInference Attack

Membership Inference Attacks for Unseen Classes

2025-06-06 · Pratiksha Thaker, Neil Kale, Zhiwei Steven Wu, Virginia Smith

Shadow model attacks are the state-of-the-art approach for membership inference attacks on machine learning models. However, these attacks typically assume an adversary has access to a background (nonmember) data distrib…

quantile regressionregression

R.R.: Unveiling LLM Training Privacy through Recollection and Ranking

2025-02-18 · Wenlong Meng, Zhenyuan Guo, Lenan Wu, Chen Gong 외

Large Language Models (LLMs) pose significant privacy risks, potentially leaking training data due to implicit memorization. Existing privacy attacks primarily focus on membership inference attacks (MIAs) or data extract…

Memorization