To share or not to share: What risks would laypeople accept to give sensitive data to differentially-private NLP systems?
Although the NLP community has adopted central differential privacy as a go-to framework for privacy-preserving model training or data sharing, the choice and interpretation of the key parameter, privacy budget $\varepsilon$ that governs the strength of privacy protection, remains largely arbitrary. We argue that determining the $\varepsilon$ value should not be solely in the hands of researchers or system developers, but must also take into account the actual people who share their potentially sensitive data. In other words: Would you share your instant messages for $\varepsilon$ of 10? We address this research gap by designing, implementing, and conducting a behavioral experiment (311 lay participants) to study the behavior of people in uncertain decision-making situations with respect to privacy-threatening situations. Framing the risk perception in terms of two realistic NLP scenarios and using a vignette behavioral study help us determine what $\varepsilon$ thresholds would lead lay people to be willing to share sensitive textual data - to our knowledge, the first study of its kind.
Code (0)
등록된 구현이 없습니다.
Tasks
Decision MakingPrivacy PreservingSimilar Papers 제목 키워드 기반
Generics in science communication: Misaligned interpretations across laypeople, scientists, and large language models
Scientists often use generics, that is, unquantified statements about whole categories of people or phenomena, when communicating research findings (e.g., "statins reduce cardiovascular events"). Large language models (L…
Gender Bias in LLMs: Preliminary Evidence from Shared Parenting Scenario in Czech Family Law
Access to justice remains limited for many people, leading laypersons to increasingly rely on Large Language Models (LLMs) for legal self-help. Laypeople use these tools intuitively, which may lead them to form expectati…
How Much Can AI Understand? Toward AI-Assisted Sensemaking of Collaborative Discussion in Groups with Shared History
AI tools that support collaborative discussion typically treat the discussion as a standalone task, focusing only on its content and setting aside the social context of the group having it. But it is groups with a shared…
Agent-Supported Foresight for AI Systemic Risks: AI Agents for Breadth, Experts for Judgment
AI impact assessments often stress near-term risks because human judgment degrades over longer horizons, exemplifying the Collingridge dilemma: foresight is most needed when knowledge is scarcest. To address long-term sy…
Asset Prices and Capital Share Risks: Theory and Evidence
An asset pricing model using long-run capital share growth risk has recently been found to successfully explain U.S. stock returns. Our paper adopts a recursive preference utility framework to derive an heterogeneous ass…