paper-with-me

홈 › Papers

When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills

2026-08-04 · Yongli Xiang, Zhifang Zhang, Bojun Yang, Ziming Hong, Lei Feng, Miao Xu, Tongliang Liu arxiv

Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personalization, this process concentrates fragmented personal signals, amplifies their impact through reuse, and challenges defenses designed for individual records or retrieval-based memory. To systematically investigate the safety of the persona-skill pipeline, we introduce AntiSkillBench, an end-to-end benchmark for evaluating risks and defenses across the persona-skill pipeline. It comprises: (i) a dataset of 7,500 persona-grounded dialogue traces, constructed from 50 behaviorally rich profiles spanning diverse task scenarios; (ii) an evaluation suite that measures skill-level privacy leakage and agent-level attribute disclosure and behavioral impersonation across three skill-distillation strategies; and (iii) a defense evaluation covering four configurations across online and post-hoc interventions, including active risk suppression and passive provenance protection. Experiments across three frontier agents show that persona-skill risks persist across agent backbones and distillation protocols, extending from explicit attributes to communication styles and personality traits. Existing defenses exhibit limited and distillation-dependent effectiveness, failing to generalize across risk and distillation strategies. These results highlight AntiSkillBench as a challenging benchmark for developing privacy-preserving and authenticity-aware persona skills.

📄 PDF Abstract BibTeX arXiv:2608.03700

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents

2026-04-23 · Wenjie Fu, Xiaoting Qin, Jue Zhang, Qingwei Lin 외 arxiv

Enterprise LLM agents can dramatically improve workplace productivity, but their core capability, retrieving and using internal context to act on a user's behalf, also creates new risks for sensitive information leakage.…

Privacy Leakage Overshadowed by Views of AI: A Study on Human Oversight of Privacy in Language Model Agent

2024-11-02 · Zhiping Zhang, Bingcan Guo, Tianshi Li

Language model (LM) agents that act on users' behalf for personal tasks (e.g., replying emails) can boost productivity, but are also susceptible to unintended privacy leakage risks. We present the first study on people's…

Language ModelingLanguage ModellingPrivacy Preserving

SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought

2025-11-11 · Shourya Batra, Pierce Tillman, Samarth Gaggar, Shashank Kesineni 외 arxiv

As Large Language Models (LLMs) evolve into personal assistants with access to sensitive user data, they face a critical privacy challenge: while prior work has addressed output-level privacy, recent findings reveal that…

AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents

2025-03-12 · Arman Zharmagambetov, Chuan Guo, Ivan Evtimov, Maya Pavlova 외

LLM-powered AI agents are an emerging frontier with tremendous potential to increase human productivity. However, empowering AI agents to take action on their user's behalf in day-to-day tasks involves giving them access…

PDSL: Privacy-Preserved Decentralized Stochastic Learning with Heterogeneous Data Distribution

2025-03-31 · Lina Wang, Yunsheng Yuan, Chunxiao Wang, Feng Li

In the paradigm of decentralized learning, a group of agents collaborates to learn a global model using distributed datasets without a central server. However, due to the heterogeneity of the local data across the differ…