paper-with-me

홈 › Papers

MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models

2025-11-21 · Xiongtao Sun, Hui Li, Jiaming Zhang, Yujie Yang, Kaili Liu, Ruxin Feng, Wen Jun Tan, Wei Yang Bryan Lim arxiv

Modern Vision-Language Models (VLMs) pose significant individual-level privacy risks by linking fragmented multimodal data to identifiable individuals through hierarchical chain-of-thought reasoning. However, existing privacy benchmarks remain structurally insufficient for this threat, as they primarily evaluate privacy perception while failing to address the more critical risk of privacy reasoning: a VLM's ability to infer and link distributed information to construct individual profiles. To address this gap, we propose MultiPriv, the first benchmark designed to systematically evaluate individual-level privacy reasoning in VLMs. We introduce the Privacy Perception and Reasoning (PPR) framework and construct a bilingual multimodal dataset with synthetic individual profiles, where identifiers, such as faces and names, are linked to sensitive attributes. This design enables nine challenging tasks spanning attribute detection, cross-image re-identification, and chained inference. We conduct a large-scale evaluation of over 50 open-source and commercial VLMs. In our controlled benchmark, 60% of widely used VLMs can perform individual-level privacy reasoning with up to 80% accuracy, suggesting a significant potential threat to personal privacy. The benchmark is available at https://github.com/CyberChangAn/MultiPriv-PII.

📄 PDF Abstract BibTeX arXiv:2511.16940

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

IDP-Bench: Benchmarking ability of LLMs to protect personal information in interdependent privacy contexts

2026-06-06 · Ayana Hussain, Soumya Sharma, Golnoosh Farnadi, Nicholas Vincent 외 arxiv

Large language models (LLMs) are becoming widely deployed as personal AI assistants with access to sensitive user data, making privacy a major challenge for their design and evaluation. Prior work focuses mainly on indiv…

DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy

2026-04-17 · Erchi Wang, Pengrun Huang, Eli Chien, Om Thakkar 외 arxiv

Differential privacy (DP) has a wide range of applications for protecting data privacy, but designing and verifying DP algorithms requires expert-level reasoning, creating a high barrier for non-expert practitioners. Pri…

Mathematical Reasoning

Beyond Verification: Abductive Explanations for Post-AI Assessment of Privacy Leakage

2025-11-13 · Belona Sonna, Alban Grastien, Claire Benn arxiv

Privacy leakage in AI-based decision processes poses significant risks, particularly when sensitive information can be inferred. We propose a formal framework to audit privacy leakage using abductive explanations, which …

FLOW: A Feedback-Driven Synthetic Longitudinal Dataset of Work and Wellbeing

2025-12-28 · Wafaa El Husseini arxiv

Access to longitudinal, individual-level data on work-life balance and wellbeing is limited by privacy, ethical, and logistical constraints. This poses challenges for reproducible research, methodological benchmarking, a…

GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs

2026-06-02 · Weidong Tang, Jierui Li, Yueling Hou, Zihan Mei 외 arxiv

True general intelligence requires not only a model of the physical world but also a social world model: the capacity to infer how individual mental states interact and crystallize into group-level outcomes. Despite nota…